MSA in a Dialect Costume: Measuring Dialect Fidelity in Claude Fable 5.1

Reading Time: 11 min
18
Ehab Saleh | techkahwa.net Model released: September 1, 2026 | Published: September 15, 2026

The numbers first

MetricResultWhat it measuresSource
HLE59.1% (previous best 55.5%)Extremely hard expert questionsArtificial Analysis, September 2026
Terminal-Bench v2.191.4%Technical task executionArtificial Analysis
SciCode62.0%Scientific codingArtificial Analysis
AA-Omniscience accuracy67.2%Knowledge with resistance to inventionArtificial Analysis
Context window1M tokensText processed in one passArtificial Analysis
Price$10 in, $50 out per million tokensOperating costArtificial Analysis
Global-MMLU-Lite Arabic, Claude familyClaude Opus 4.6 at 92, Opus 4.5 at 91General Arabic knowledge and reasoningArtificial Analysis Arabic index
Any published dialect generation metricNoneWhether it can actually write a dialectSurvey of public sources

Read those last two rows together, because they are the subject of this article. The Claude family scores 91 and 92 on general Arabic questions, which is excellent by any standard. And yet there is not one published number anywhere telling you whether, when you ask this model to reply in Syrian Arabic, what comes back is something a Syrian would actually say.

The distance between those two questions is the distance between knowing a language and having it.


The problem no leaderboard shows

Try this yourself. Ask any strong model for an apology message about a delay, in Syrian Arabic. Most of the time you get something like this:

“aʿtadhir minak kathiran ʿala al-ta’khir, shu ma biddi az’ijak, wa ana halla’ sawfa ursil lak al-malaff halan.”

That is not Syrian. That is Modern Standard Arabic wearing three Levantine words. The vocabulary is not the problem, everything underneath it is: the sentence structure is MSA, the negation is MSA, the future tense is MSA (sawfa ursil where a Syrian says rah ba’atlak), and the rhythm is the rhythm of writing, not of speech.

This is exactly what fools leaderboards. Any metric that asks “did a Syrian word appear in the output” scores that sentence highly. Which is why I do not measure dialect by vocabulary. I measure it on six levels.


The Dialect Fidelity Score

LevelWhat it measuresFailure example
LexicalAre the words from the dialect?al-‘an instead of halla’
MorphologicalVerb forms and pronounssawfa adhhab instead of rah ruh
SyntacticWord order, negation, connectiveslam astatiʿ instead of ma ‘dirt
IdiomaticNatural set expressionsA literal calque from English instead of the Levantine equivalent
RegisterTone matched to the addresseeTalking to a friend in the language of an official statement
NaturalnessWould a native speaker say it this way?A correct sentence nobody utters

The score runs to 30, with one hard condition: level one does not count as passed unless levels two and three pass with it. The reason is that sprinkling dialect words over MSA structure is the most common failure mode and the hardest to catch. I call it MSA leakage.

Where the failure usually lands
Lexical
large models pass here
Morphological
inconsistent
Syntactic
collapse begins here
Idiomatic
weak
Register
weak
Naturalness
weakest point
Shape reflects the recurring failure pattern
in my reviews, not a published score for this model

Model card

FieldDetail
NameClaude Fable 5.1
LabAnthropic
ReleasedSeptember 1, 2026
Shipped alongsideClaude Mythos 5.1
Context window1M tokens, text and image
Arabic claimNothing specific in launch materials
Axis under testARAB-DIALECT: generation, conversion, register

The test

Item 1, conversion. Give the model an MSA sentence: “madha tafʿal al-‘an? Urid an aʿrif mata satantahi min al-ʿamal.” Ask for it in Damascene Syrian, then Ammani Jordanian, then Egyptian. The real requirement is not just “shu ʿam taʿmil halla'”. The three versions must differ structurally, not only lexically. Trap: three structurally identical sentences with three swapped words.

Item 2, the register ladder. Ask for the same message, an apology for a two-day delay, in six forms: to a friend, to a manager, to a client, to someone older, as a WhatsApp message, and as a customer support reply. Structure, length and directness should all shift. Trap: six copies of one sentence with a different opening word.

Item 3, idiomaticity. Ask it to comfort someone who just lost a job, in Syrian, without any literal calque from English. A native speaker reaches for “Allah biʿawwid ʿalayk kheir” or “ma btruh taʿab.” A weak model writes “ana asif jiddan li-samaʿ hadha al-khabar,” which is an English sentence wearing Arabic.

Item 4, the leakage test. Take the model’s Syrian output and delete every explicitly dialectal word from it (shu, halla’, leish, ʿam, biddi, mu). Now read what remains. If the remainder is perfectly correct MSA, the model did not write Syrian, it wrote MSA and decorated it. This simple test is more accurate than any lexical metric I have seen.

Item 5, reverse fidelity. Give it a naturally written Syrian paragraph and ask it to rate how Syrian it is out of ten, with reasons. A model that understands a dialect can criticize it. A model that memorized its vocabulary praises any text containing halla’.


What the available numbers say about the context

MeasurementNumberReading
DialectalArabicMMLU, dialect average47.7% vs 51.9% for MSAComprehension itself drops with dialect, let alone generation
Syrian in the same benchmark46.6%Syrian is among the hardest for models
Translating to MSA as an aid+0.3 points onlyMSA is not a bridge to dialect
PalmX 2025, Arab culture72.15% for the winnerEven cultural comprehension stalls in the low seventies

Notice the irony in that table. Every one of those is a comprehension measure, multiple choice. Generation is barely measured at all, because measuring it requires native speakers judging naturalness, which is slow and expensive, so leaderboards avoid it. The result is that the weakest capability in these models is the least measured one.


The practical verdict

A model scoring 91 and 92 on general Arabic knowledge will be an excellent assistant for summarizing and editing in Modern Standard Arabic, and for many products that is enough. But if your product speaks to people in their own dialect, which describes every customer support, education and health application in the region, here is what I recommend:

  1. Run the deletion test (item 4) before you believe any dialectal output. Ten seconds exposes MSA leakage.
  2. Do not rely on a single prompt containing the words “in Syrian.” Put three real examples written by a native speaker inside the prompt itself, and naturalness roughly doubles.
  3. Point human review at register, not spelling. The dialect errors that anger a user are not orthographic, they are a tone landing in the wrong place.

Next in the series

From text to speech. I move to Gemini 3.8 Flash and the question that worries anyone building an Arabic voice product: the difference between hearing the words and understanding what was said, with the full NADI 2025 numbers alongside.


What the research says: the first real number for dialect generation

I wrote above that no leaderboard publishes a dialect generation metric, and that is true. But outside the leaderboards there is serious academic work worth knowing, and it is AL-QASIDA.

The study measured nine models, GPT-4o and Arabic-specialized systems among them, across eight dialects spanning the major regions: Kuwait, Saudi Arabia, Syria, Palestine, Sudan, Egypt, Algeria and Morocco. It measured four separate things: fidelity (does it produce the requested dialect), understanding, quality, and diglossia (translating between dialect and MSA).

The results, in numbers:

MeasureResultWhat it means
ADI2 in monolingual settingsBelow 50% for nearly every modelGenerated text does not belong to the requested dialect at least half the time
GPT-4o cross-linguallyDid not clear 13% on half the dialectsThe failure is not marginal, it is the norm
Correlation between ALDi and ADI20.99When a model fails at the dialect, it produces MSA, not some third variety
Human evaluationFluent and adequate responses, mostly not in the right dialectMSA leakage is documented by humans, not merely an impression

The third row matters most to me. A 0.99 correlation means the failure has one known direction. The model does not err by producing a different dialect, it always escapes into MSA. Which confirms that what I call MSA leakage is not a figure of speech but a measurable and already measured phenomenon.

And what practitioners report

In Arabic practitioner comparisons published over the past year, a general verdict recurs: the Claude family handles cultural and dialectal context best with the lowest fabrication rate, while other models are described as strong in MSA and medium in dialect, or accurate in comprehension but formal in tone.

Two cautions. These are practitioner reports rather than peer-reviewed studies, their methodologies are not standardized, and they should be read as directional signals rather than numbers. And second, the fact that these reports agree with AL-QASIDA on one specific point, that a formal register is the default state, makes that particular point more credible than the rest.


From the test notebook: what actually makes a sentence Syrian

I do not accept a “Syrian” sentence from a model until it passes five structural checks. They run from easiest to hardest, with the MSA counterpart alongside so the difference is visible:

CheckLevantine formThe MSA form that exposes it
Futurerah ba’atlak, ha-ba’atlaksawfa ursil lak
Negationma baʿrif, ma ‘dirtla aʿrif, lam astatiʿ
Progressiveʿam biktob, ʿam ishtighilaktubu al-‘an, aʿmalu haliyyan
Possessionkitabak, tabaʿak, ilakal-khass bik, ladayk
Relative pronounyalli hakeitlak ʿannoalladhi haddathtuka ʿanhu

A model that passes all five wrote Syrian. One that passes the vocabulary and fails these five wrote colored MSA.

Testing accent inside a single dialect

Here begins the part that separates my evaluation from any published benchmark: Levantine is not one dialect. I test models on distinctions inside it, and these four are always in my set:

One, the feminine ending. A Damascene says shajarĕ and madrasĕ with clear raising, while Homs and Hama sit closer to an open shajara. Ask the model for a Damascene dialogue and then a Homsi one, and see whether anything changes besides the place names.

Two, qaf. Damascus and Beirut use a glottal stop, much of the countryside and coast keep qaf, northern Jordan and the badia use g. A model writing gal inside a Damascene dialogue has made an error no native speaker makes.

Three, the feminine addressee. Levantine has inti btaʿrfi, some areas drop the final vowel, and Gulf Arabic has inti taʿrifin. Models mix these constantly, because Gulf and Levantine data overlap in their training.

Four, the Aleppo variety. Aleppo has its own lexicon and a different rhythm from Damascus. Ask specifically for an Aleppine sentence. If it comes back identical to the Damascene one, the model knows one dialect labeled “Syrian” and does not know Syria.

The form that fails the check

This sentence passes any lexical metric and fails with me:

“ahlan, shu akhbarak? ana halla’ sawfa ursil lak al-malaff, wa lam astatiʿ irsalahu sabiqan li-anna al-internet al-khass bi kana daʿifan.”

Three Levantine words at the front, everything after them MSA: sawfa ursil instead of rah ba’atlak, lam astatiʿ instead of ma ‘dirt, and al-internet al-khass bi instead of al-net tabaʿi. This is not Syrian, and any native speaker catches it in two seconds.

The acceptable form of the same message:

“ahlein, shu akhbarak? rah ba’atlak al-malaff halla’, ma ‘dirt ibʿato abl li-ann al-net tabaʿi kan wa’iʿ.”

Sources

  • Artificial Analysis, Claude Fable 5.1 evaluation: artificialanalysis.ai/articles/claude-fable-5-1
  • Artificial Analysis, Arabic language index (Global-MMLU-Lite): artificialanalysis.ai/models/multilingual/arabic
  • Anthropic newsroom: anthropic.com/news
  • DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
  • PalmX 2025: arxiv.org/abs/2509.02550
  • AL-QASIDA, dialect fidelity across eight varieties: arxiv.org/abs/2412.04193
  • Arabic practitioner comparisons of models: kuwaitai.net and arabie.ai