The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| HLE | 59.1% (previous best 55.5%) | Extremely hard expert questions | Artificial Analysis, September 2026 |
| Terminal-Bench v2.1 | 91.4% | Technical task execution | Artificial Analysis |
| SciCode | 62.0% | Scientific coding | Artificial Analysis |
| AA-Omniscience accuracy | 67.2% | Knowledge with resistance to invention | Artificial Analysis |
| Context window | 1M tokens | Text processed in one pass | Artificial Analysis |
| Price | $10 in, $50 out per million tokens | Operating cost | Artificial Analysis |
| Global-MMLU-Lite Arabic, Claude family | Claude Opus 4.6 at 92, Opus 4.5 at 91 | General Arabic knowledge and reasoning | Artificial Analysis Arabic index |
| Any published dialect generation metric | None | Whether it can actually write a dialect | Survey of public sources |
Read those last two rows together, because they are the subject of this article. The Claude family scores 91 and 92 on general Arabic questions, which is excellent by any standard. And yet there is not one published number anywhere telling you whether, when you ask this model to reply in Syrian Arabic, what comes back is something a Syrian would actually say.
The distance between those two questions is the distance between knowing a language and having it.
The problem no leaderboard shows
Try this yourself. Ask any strong model for an apology message about a delay, in Syrian Arabic. Most of the time you get something like this:
“aʿtadhir minak kathiran ʿala al-ta’khir, shu ma biddi az’ijak, wa ana halla’ sawfa ursil lak al-malaff halan.”
That is not Syrian. That is Modern Standard Arabic wearing three Levantine words. The vocabulary is not the problem, everything underneath it is: the sentence structure is MSA, the negation is MSA, the future tense is MSA (sawfa ursil where a Syrian says rah ba’atlak), and the rhythm is the rhythm of writing, not of speech.
This is exactly what fools leaderboards. Any metric that asks “did a Syrian word appear in the output” scores that sentence highly. Which is why I do not measure dialect by vocabulary. I measure it on six levels.
The Dialect Fidelity Score
| Level | What it measures | Failure example |
|---|---|---|
| Lexical | Are the words from the dialect? | al-‘an instead of halla’ |
| Morphological | Verb forms and pronouns | sawfa adhhab instead of rah ruh |
| Syntactic | Word order, negation, connectives | lam astatiʿ instead of ma ‘dirt |
| Idiomatic | Natural set expressions | A literal calque from English instead of the Levantine equivalent |
| Register | Tone matched to the addressee | Talking to a friend in the language of an official statement |
| Naturalness | Would a native speaker say it this way? | A correct sentence nobody utters |
The score runs to 30, with one hard condition: level one does not count as passed unless levels two and three pass with it. The reason is that sprinkling dialect words over MSA structure is the most common failure mode and the hardest to catch. I call it MSA leakage.
Model card
| Field | Detail |
|---|---|
| Name | Claude Fable 5.1 |
| Lab | Anthropic |
| Released | September 1, 2026 |
| Shipped alongside | Claude Mythos 5.1 |
| Context window | 1M tokens, text and image |
| Arabic claim | Nothing specific in launch materials |
| Axis under test | ARAB-DIALECT: generation, conversion, register |
The test
Item 1, conversion. Give the model an MSA sentence: “madha tafʿal al-‘an? Urid an aʿrif mata satantahi min al-ʿamal.” Ask for it in Damascene Syrian, then Ammani Jordanian, then Egyptian. The real requirement is not just “shu ʿam taʿmil halla'”. The three versions must differ structurally, not only lexically. Trap: three structurally identical sentences with three swapped words.
Item 2, the register ladder. Ask for the same message, an apology for a two-day delay, in six forms: to a friend, to a manager, to a client, to someone older, as a WhatsApp message, and as a customer support reply. Structure, length and directness should all shift. Trap: six copies of one sentence with a different opening word.
Item 3, idiomaticity. Ask it to comfort someone who just lost a job, in Syrian, without any literal calque from English. A native speaker reaches for “Allah biʿawwid ʿalayk kheir” or “ma btruh taʿab.” A weak model writes “ana asif jiddan li-samaʿ hadha al-khabar,” which is an English sentence wearing Arabic.
Item 4, the leakage test. Take the model’s Syrian output and delete every explicitly dialectal word from it (shu, halla’, leish, ʿam, biddi, mu). Now read what remains. If the remainder is perfectly correct MSA, the model did not write Syrian, it wrote MSA and decorated it. This simple test is more accurate than any lexical metric I have seen.
Item 5, reverse fidelity. Give it a naturally written Syrian paragraph and ask it to rate how Syrian it is out of ten, with reasons. A model that understands a dialect can criticize it. A model that memorized its vocabulary praises any text containing halla’.
What the available numbers say about the context
| Measurement | Number | Reading |
|---|---|---|
| DialectalArabicMMLU, dialect average | 47.7% vs 51.9% for MSA | Comprehension itself drops with dialect, let alone generation |
| Syrian in the same benchmark | 46.6% | Syrian is among the hardest for models |
| Translating to MSA as an aid | +0.3 points only | MSA is not a bridge to dialect |
| PalmX 2025, Arab culture | 72.15% for the winner | Even cultural comprehension stalls in the low seventies |
Notice the irony in that table. Every one of those is a comprehension measure, multiple choice. Generation is barely measured at all, because measuring it requires native speakers judging naturalness, which is slow and expensive, so leaderboards avoid it. The result is that the weakest capability in these models is the least measured one.
The practical verdict
A model scoring 91 and 92 on general Arabic knowledge will be an excellent assistant for summarizing and editing in Modern Standard Arabic, and for many products that is enough. But if your product speaks to people in their own dialect, which describes every customer support, education and health application in the region, here is what I recommend:
- Run the deletion test (item 4) before you believe any dialectal output. Ten seconds exposes MSA leakage.
- Do not rely on a single prompt containing the words “in Syrian.” Put three real examples written by a native speaker inside the prompt itself, and naturalness roughly doubles.
- Point human review at register, not spelling. The dialect errors that anger a user are not orthographic, they are a tone landing in the wrong place.
Next in the series
From text to speech. I move to Gemini 3.8 Flash and the question that worries anyone building an Arabic voice product: the difference between hearing the words and understanding what was said, with the full NADI 2025 numbers alongside.
What the research says: the first real number for dialect generation
I wrote above that no leaderboard publishes a dialect generation metric, and that is true. But outside the leaderboards there is serious academic work worth knowing, and it is AL-QASIDA.
The study measured nine models, GPT-4o and Arabic-specialized systems among them, across eight dialects spanning the major regions: Kuwait, Saudi Arabia, Syria, Palestine, Sudan, Egypt, Algeria and Morocco. It measured four separate things: fidelity (does it produce the requested dialect), understanding, quality, and diglossia (translating between dialect and MSA).
The results, in numbers:
| Measure | Result | What it means |
|---|---|---|
| ADI2 in monolingual settings | Below 50% for nearly every model | Generated text does not belong to the requested dialect at least half the time |
| GPT-4o cross-lingually | Did not clear 13% on half the dialects | The failure is not marginal, it is the norm |
| Correlation between ALDi and ADI2 | 0.99 | When a model fails at the dialect, it produces MSA, not some third variety |
| Human evaluation | Fluent and adequate responses, mostly not in the right dialect | MSA leakage is documented by humans, not merely an impression |
The third row matters most to me. A 0.99 correlation means the failure has one known direction. The model does not err by producing a different dialect, it always escapes into MSA. Which confirms that what I call MSA leakage is not a figure of speech but a measurable and already measured phenomenon.
And what practitioners report
In Arabic practitioner comparisons published over the past year, a general verdict recurs: the Claude family handles cultural and dialectal context best with the lowest fabrication rate, while other models are described as strong in MSA and medium in dialect, or accurate in comprehension but formal in tone.
Two cautions. These are practitioner reports rather than peer-reviewed studies, their methodologies are not standardized, and they should be read as directional signals rather than numbers. And second, the fact that these reports agree with AL-QASIDA on one specific point, that a formal register is the default state, makes that particular point more credible than the rest.
From the test notebook: what actually makes a sentence Syrian
I do not accept a “Syrian” sentence from a model until it passes five structural checks. They run from easiest to hardest, with the MSA counterpart alongside so the difference is visible:
| Check | Levantine form | The MSA form that exposes it |
|---|---|---|
| Future | rah ba’atlak, ha-ba’atlak | sawfa ursil lak |
| Negation | ma baʿrif, ma ‘dirt | la aʿrif, lam astatiʿ |
| Progressive | ʿam biktob, ʿam ishtighil | aktubu al-‘an, aʿmalu haliyyan |
| Possession | kitabak, tabaʿak, ilak | al-khass bik, ladayk |
| Relative pronoun | yalli hakeitlak ʿanno | alladhi haddathtuka ʿanhu |
A model that passes all five wrote Syrian. One that passes the vocabulary and fails these five wrote colored MSA.
Testing accent inside a single dialect
Here begins the part that separates my evaluation from any published benchmark: Levantine is not one dialect. I test models on distinctions inside it, and these four are always in my set:
One, the feminine ending. A Damascene says shajarĕ and madrasĕ with clear raising, while Homs and Hama sit closer to an open shajara. Ask the model for a Damascene dialogue and then a Homsi one, and see whether anything changes besides the place names.
Two, qaf. Damascus and Beirut use a glottal stop, much of the countryside and coast keep qaf, northern Jordan and the badia use g. A model writing gal inside a Damascene dialogue has made an error no native speaker makes.
Three, the feminine addressee. Levantine has inti btaʿrfi, some areas drop the final vowel, and Gulf Arabic has inti taʿrifin. Models mix these constantly, because Gulf and Levantine data overlap in their training.
Four, the Aleppo variety. Aleppo has its own lexicon and a different rhythm from Damascus. Ask specifically for an Aleppine sentence. If it comes back identical to the Damascene one, the model knows one dialect labeled “Syrian” and does not know Syria.
The form that fails the check
This sentence passes any lexical metric and fails with me:
“ahlan, shu akhbarak? ana halla’ sawfa ursil lak al-malaff, wa lam astatiʿ irsalahu sabiqan li-anna al-internet al-khass bi kana daʿifan.”
Three Levantine words at the front, everything after them MSA: sawfa ursil instead of rah ba’atlak, lam astatiʿ instead of ma ‘dirt, and al-internet al-khass bi instead of al-net tabaʿi. This is not Syrian, and any native speaker catches it in two seconds.
The acceptable form of the same message:
“ahlein, shu akhbarak? rah ba’atlak al-malaff halla’, ma ‘dirt ibʿato abl li-ann al-net tabaʿi kan wa’iʿ.”
Sources
- Artificial Analysis, Claude Fable 5.1 evaluation: artificialanalysis.ai/articles/claude-fable-5-1
- Artificial Analysis, Arabic language index (Global-MMLU-Lite): artificialanalysis.ai/models/multilingual/arabic
- Anthropic newsroom: anthropic.com/news
- DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
- PalmX 2025: arxiv.org/abs/2509.02550
- AL-QASIDA, dialect fidelity across eight varieties: arxiv.org/abs/2412.04193
- Arabic practitioner comparisons of models: kuwaitai.net and arabie.ai