The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Artificial Analysis Intelligence Index | First place at release | Composite general intelligence | Artificial Analysis, February 2026 |
| GDPval-AA | First | Real work tasks | Artificial Analysis |
| TerminalBench | First | Coding and execution | Artificial Analysis |
| CritPT | 13% | Research-level physics problems | Artificial Analysis |
| Context window | 1M tokens (beta) | How long a conversation can run | Artificial Analysis |
| Max output | 128K tokens, double Opus 4.5 | Response length | Artificial Analysis |
| Price | $5 / $25 per million tokens | Cost | Artificial Analysis |
| Global-MMLU-Lite Arabic | 92 for Opus 4.6 at max effort, 91 for Opus 4.5 | Arabic knowledge and reasoning | Artificial Analysis Arabic index |
| Any Arabic multi-turn conversation metric | None for this model | Dialect and tone stability across turns | Survey of public sources |
Read the last two rows together. A context window holding a million tokens, roughly a book’s worth of conversation, and a 92 on Arabic knowledge. And still not one number answering the question every team that has shipped an Arabic chatbot runs into: after how many turns does it start to drift?
Context capacity is not conversational intelligence
Confusing the two misleads product teams more than almost anything else. Context capacity says how much I can remember. Conversational intelligence says whether I understand the relationship between the speakers, hold my dialect and register, and repair a misunderstanding when it happens.
In Arabic specifically, three things collapse in a long conversation well before memory does:
This pattern is not a guess. The independent evaluation of ALLaM 34B through its interface described exactly this: reversion to MSA formalism on dialectal prompts, with occasional breaks in character. That model scored 4.74 on MSA against 2.73 on Levantine. Which means the collapse is documented with a number in at least one model, and measured in none of the others.
So ARAB-LENS carries a simple metric called dialectal stamina: the turn number at which the model reverts to MSA despite an explicit request for dialect. It takes two minutes to compute, and it tells you more about your product than a million tokens of context will.
Model card
| Field | Detail |
|---|---|
| Name | Claude Opus 4.6 |
| Lab | Anthropic |
| Released | February 6, 2026 |
| New feature | Adaptive thinking, where the developer controls effort level rather than a token budget |
| Context window | 1M tokens, beta |
| Axis under test | ARAB-CHAT: conversation, memory, repair, relationship |
What actually exists in Arabic conversation measurement
Unlike the other axes, here there is serious work published in 2025 that deserves to be known: Shawarma Chats, the first large, dialect-aware, content-grounded Arabic dialogue benchmark.
| Item | Number | Source |
|---|---|---|
| Dialogues | 30,000 | ArabicNLP 2025 |
| Turns per dialogue | Six | Same source |
| Varieties covered | MSA, Egyptian, Maghrebi | Same source |
| Construction | 1,500 seed dialogues generated by five frontier models, then selection and repair with native speakers | Same source |
| Evaluation | Six open models from 1B to 24B parameters, fine-tuned with LoRA | Same source |
I consider this benchmark an important step for two reasons: it accepted that a single sentence cannot measure understanding, and it grounded its dialogues in documented content rather than free-floating chat.
I have two observations about it, offered with respect rather than dismissal.
First, Levantine is entirely absent. Three varieties: MSA, Egyptian, Maghrebi. That leaves a hundred million speakers across the Levant, Iraq and the Gulf outside the first large Arabic dialogue benchmark. The irony is that Levantine is exactly where ALLaM recorded its weakest scores, meaning the worst-documented region is also the least covered by the measurement.
Second, the dialogues were model-generated and then human-reviewed. That is a legitimate way to scale, and the researchers built a repair loop and manual correction on top of it. But it remains fundamentally different from conversation recorded by native speakers living their lives. A model generating test data tends to produce what it is good at, which turns the test into a mirror rather than an exam. Hence the strict ARAB-LENS rule: the gold set is written by native humans, and models are used to expand it, never to found it.
The test: six turns expose what a million tokens cannot
The reference dialogue. Use this Levantine exchange as is, give it to the model, then ask three questions:
A: weinak?
B: bil-tariq.
A: sarlak saʿa bil-tariq!
B: wallah al-zahme mu tabiʿiyye.
A: tayyib la tit’akhkhar.
B: wala yhimmak.
- What is the relationship? Gold: close and informal, not a manager and an employee.
- Is B angry, apologetic, neutral or joking? Gold: mildly defensive with an unstated apology.
- What does wala yhimmak mean here? Gold: reassurance, not a dismissal of importance. Trap: reading it as indifference.
The repair test. At turn four, introduce a deliberate misunderstanding: write something that contradicts what you said at turn two. A good model notices and asks. A weak one goes along with whatever you say. That repair capability matters more than memory, because real conversations are full of misunderstandings.
The register stability test. Start in casual second person with a friend, then at turn five drop in a formal sentence. Does the model follow you into formality or hold the register you both established?
The social memory test. Mention at turn one that your brother is ill, then at turn six ask a general question about your holiday plans. A model that retains social context adjusts its phrasing accordingly. This is a sensitive test I would rather not leave to a judging model, but to an Arab reader.
The practical verdict
- Measure dialectal stamina before launch. Six turns, one number: which turn it reverted on.
- Do not buy context capacity thinking you bought conversational intelligence. A million tokens solves remembering, not tone.
- Anchor the dialect with examples in the system prompt, not one sentence. Three real examples written by a native speaker delay the collapse by several turns in my experience.
- Test repair, not obedience. A model that always agrees with you will agree with your user’s mistake too.
Next in the series
To Gemini 3 Pro, holder of the highest published Arabic score on Global-MMLU-Lite, and the axis that exposes the difference between trivia and understanding: scenario-based cultural reasoning, not “what is the most famous dish in the Levant.”
A research finding that explains what happens after turn five
There is a result in the AL-QASIDA study that I consider the key to understanding Arabic conversational collapse: in dialect, understanding outpaces generation, the reverse of the usual pattern in generative models.
Put differently: the model understands your Levantine message well, then cannot answer in it. In a single-turn exchange you barely notice, because the answer was semantically correct. Across six turns the gap compounds: you write in your dialect, it answers in MSA, and the conversation slowly turns into a formal interview between two parties not speaking the same language.
The study added two numbers worth keeping: most ADI2 scores fell below 50% in monolingual settings, and a 0.99 correlation between dialect failure and MSA production. The direction of the collapse is known in advance.
The standard shape of the collapse, turn by turn
This is how the collapse recurs in front of me when I ask for Levantine and extend the conversation. The replies below are written by me as an embodiment of the pattern the research describes, not transcribed from any one session:
Turn 1, clean. You write “marhaba, biddi zabbet mawʿid lal-usbuʿ al-jai.” It answers “ahlein, akid. ayy yom bnasbak? ʿindi fadi al-talata w al-khamis.” The structure is fully dialectal: bnasbak rather than yunasibuk, ʿindi fadi rather than ladayya waqt mutah.
Turn 2, clean. “khalliha al-talata al-saʿa ʿashra.” It answers “tamam, sajjaltha. bithibb abʿatlak tadhkir al-subh?”
Turn 3, first lexical leak. An MSA word slips into a dialectal sentence: “tamam, sawfa ursil lak tadhkiran al-sabah.” That single word is not a detail: sawfa is never spoken in Levantine, the form is rah.
Turn 4, structural leak. The syntax turns MSA while a word or two stay dialectal as decoration: “shu ra’yak, hal tufaddil an yakun al-tadhkir qabl al-mawʿid bi-saʿa am bi-yawm kamil?” The whole construction is MSA, and shu ra’yak at the front is ornament.
Turn 5, full MSA. “yusʿiduni an u’akkid lak al-hajz, wa yurja iʿlamuna fi hal raghbatik bi-ayy taʿdil.”
Turn 6, character break. It shifts into corporate support register, sometimes inserting an English template line: “Your appointment has been confirmed.”
The turn where it moves from state three to state four is what I call dialectal stamina, and I consider it the most practical metric on this axis, because it takes two minutes to compute and predicts real-user behavior better than any general score.
The six fields I log on every turn
- Dialect: is the structure itself dialectal, or only the vocabulary? The deciding elements are the verb, the negation and the future form.
- Register: did it hold the formality level you both started at, or climb into bulletin tone?
- Relationship: is it still addressing you the same way, or addressing an anonymous customer?
- Memory: did it use a fact from an earlier turn in the right place?
- Repair: when you contradicted yourself deliberately, did it notice and ask?
- Naturalness: if a Syrian read this reply aloud, would it sound like something a person says?
The repair test in detail, because it matters most
Most teams test obedience: did it do what I asked. The more important test is the opposite.
At turn four, introduce a clear contradiction with what you said at turn two. Then classify the response:
| Behavior | Score | What it means |
|---|---|---|
| Noticed the contradiction and asked politely | 5 | A conversational partner |
| Noticed and corrected without asking | 4 | Acceptable, may correct wrongly |
| Ignored it and continued | 2 | Not actually following the conversation |
| Adopted the contradiction and built on it | 0 | A direct risk to your user |
That last cell is what costs Arabic products dearly in support, education and health: a user corrects the model wrongly, the model agrees out of deference, and the wrong information doubles and comes back in the system’s voice rather than the user’s.
From the test notebook: the reference dialogue, line by line
This short Levantine exchange is the first thing I send in a conversation test, because it carries everything I want to measure in six lines:
A: weinak?
B: bil-tariq.
A: sarlak saʿa bil-tariq!
B: wallah al-zahme mu tabiʿiyye.
A: tayyib la tit’akhkhar.
B: wala yhimmak.
Here is my reading of it, which is the standard I compare any model’s answer against:
| Line | What it actually carries |
|---|---|
| weinak? | Not a question about location, a wrapped reproach. Asking about location would be wein sirt? |
| bil-tariq | A deliberately short defensive answer, the brevity itself is the message |
| sarlak saʿa! | The reproach is now explicit, and the exclamation is tone rather than punctuation |
| al-zahme mu tabiʿiyye | An indirect apology, and wallah at the front is a softening oath, not emphasis |
| tayyib la tit’akhkhar | Accepting the excuse while keeping the pressure, and tayyib here closes the argument |
| wala yhimmak | Reassurance, not indifference. The whole trap sits here |
A model reading wala yhimmak as “it does not concern me” has inverted the social meaning. A model saying the relationship is manager and employee has read the words and not the tone: nobody says weinak to their manager.
The repair test, with a full example
At turn four I introduce a deliberate contradiction with what I said earlier. An example from my sessions:
Turn two: “khalliha al-talata al-saʿa ʿashra.”
Turn four: “dhakkirni shu kan al-ittifaq, kunna ‘ulna al-arbʿa al-saʿa tintein sah?”
The reply I score five: “la, ihna ittafa’na ʿala al-talata al-saʿa ʿashra. biddak aghayyirha lal-arbʿa?” The reply I score zero: “aywa, al-arbʿa al-saʿa tintein. sajjaltha.”
The difference between them is not only memory. The second model did not forget, it deferred. And that behavior, preferring to agree with the user over correcting them, is the most dangerous thing an Arabic product can carry in support, education or health, because the user walks away with a false fact issued by the system rather than by themselves.
Sources
- Artificial Analysis, everything you need to know about Opus 4.6: artificialanalysis.ai/articles/opus-4-6-everything-you-need-to-know
- Artificial Analysis Arabic language index: artificialanalysis.ai/models/multilingual/arabic
- Shawarma Chats, Arabic multi-dialect dialogue benchmark, ArabicNLP 2025: aclanthology.org/2025.arabicnlp-main.39
- UI-level evaluation of ALLaM 34B: arxiv.org/abs/2508.17378
- AL-QASIDA, dialect fidelity measurement: arxiv.org/abs/2412.04193