The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| HELM Arabic | Llama 4 Maverick among the top ten open-weights models | Seven native Arabic benchmarks | Stanford CRFM and Arabic.AI, December 2025 |
| OALL leaderboard, chat models | Llama3.3-70B-Instruct in first place | Native Arabic benchmarks | OALL v2 |
| Llama 4 Maverick | 17B active of 400B total, 1M token context | Size and context window | Wikipedia, citing Meta |
| Llama 4 Scout | 17B active of 109B total, 10M token context | Size and context window | Same source |
| Officially supported languages | 12 | Coverage breadth | Same source |
| License | Source-available with use restrictions, not fully open | Limits on use and redistribution | Same source |
| Any published MSA leakage metric | None | MSA structure surviving inside dialectal text | Survey of public sources |
Twelve languages. Arabic is one of them. That decision is worth noticing on its own: when a global lab picks only twelve languages and puts Arabic among them, it is acknowledging the language’s weight. My question is: which Arabic did they pick?
The clearest failure pattern in Arabic, and the least measured
Ask a strong model for a reply in Syrian Arabic and it will comply. That is exactly where the problem starts, because its compliance splits into two layers: a lexical layer where it obeys you, and a structural layer where it does not.
This is what I call MSA leakage. The model here is not disobeying instructions, it is obeying them literally: you asked for Syrian words and it gave you Syrian words. The problem is that a dialect is not a word list, it is a complete linguistic system.
Why does nobody catch it? Because catching it requires a native speaker to read the text and say “this is not Syrian,” and that judgment cannot be automated with multiple choice. So the most common failure mode in Arabic products is the least visible one on leaderboards.
There is at least one numerical trace of it: the independent evaluation of ALLaM 34B described the pattern explicitly, reversion to MSA formalism on dialectal prompts, and recorded 4.74 for MSA against 2.73 for Levantine in the same model.
Model card
| Field | Detail |
|---|---|
| Name | Llama 4, Scout and Maverick variants |
| Lab | Meta |
| Released | April 5, 2025 |
| Architecture | Mixture of experts, text and image input |
| Context window | 10M tokens on Scout, 1M on Maverick |
| Axis under test | ARAB-DIALECT: MSA leakage |
That ten million token window is a remarkable number, enough to analyze an entire archive of Arabic conversations in one pass. To me that is the most practical opportunity this model offers Arab teams: not writing you good dialect, but reading ten thousand of your users’ messages and classifying them.
The deletion test: ten seconds that expose leakage
This is the simplest test in all of ARAB-LENS and the most useful, and it needs no tooling and no account:
- Ask for a reply in your dialect.
- Delete every explicitly dialectal word from the reply: shu, halla’, leish, ʿam, biddi, mu, heik, la-hon.
- Read what remains.
If what remains is perfectly correct MSA, the model did not write a dialect. It wrote MSA and decorated it.
The control: take a paragraph written by an actual Syrian and delete the same words. What remains is not clean MSA, it is broken text, because the dialect is inside the structure rather than sprinkled on top of it.
Five additional items
Item 1, future tense. Ask for a sentence about tomorrow. Rah or ha against sawfa or sa. The fastest structural detector in Levantine and Egyptian.
Item 2, negation. Ask it to negate a sentence. Ma baʿrif against la aʿrif. Negation is among the last things models learn, because it is structural rather than lexical.
Item 3, possessives. Kitabak, ʿindak, ilak. A leaking model drifts toward al-khass bik even inside dialect, a construction nobody uses in speech.
Item 4, rhythm. Ask for three short consecutive sentences as they would be said in a voice note. A leaking model writes one long compound sentence, because that is the rhythm of writing.
Item 5, stability across turns. Five turns in dialect, and log the turn number where it reverts to MSA. That is dialectal stamina.
Why leakage happens, technically
Three overlapping causes, all leading to the same result:
- Data. Written Arabic online is mostly MSA: journalism, books, official sites. Dialect lives on social platforms, which is messier, harder to collect, and more likely to be filtered out during cleaning.
- Response tuning. Models are tuned to be “polite and professional,” and the calibration is usually English. In Arabic, “professional” translates automatically into MSA, so politeness and MSA become the same thing in the model’s behavior.
- Tokenization. A dialectal Arabic word splits into more tokens than its MSA counterpart, because MSA is far more frequent in training data. The model structurally prefers the cheaper path.
The practical consequence is that asking is not enough. The best result I have found in actual use comes from placing three real examples written by a native speaker inside the system prompt, rather than describing the dialect in words.
The practical verdict
- Run the deletion test on every dialectal output before launch. Ten seconds, no tools.
- Teach by example, not by description. Three authentic texts inside the prompt are worth ten sentences of instruction.
- Separate “professional” from “MSA” in your instructions. A reply can be highly polite and in dialect at once, which is what a good support agent in Amman does every day.
- Use the huge context window for analysis rather than generation. That is where this model genuinely wins.
Next in the series
DeepSeek R1 and a question I have found no published answer to: when a model shows its chain of thought, which language does it think in before writing to you in Arabic? And does its answer change if you force it to think in Arabic?
The numerical evidence that leakage has one direction
I described MSA leakage above as a failure pattern and said nobody measures it. Precision requires correcting that sentence: the leaderboards do not measure it, but academic research did, with a decisive result.
The AL-QASIDA study measured nine models across eight dialects and found a 0.99 correlation between a model failing to produce the dialect and it producing MSA.
That number deserves to be understood properly. A near-perfect correlation means the failure is not random:
Most ADI2 scores fell below 50% in monolingual settings, and GPT-4o did not clear 13% on half the dialects cross-lingually. The study’s human evaluation described outputs as fluent and adequate in meaning and mostly not in the requested dialect.
So what I call MSA leakage is not an impression. It is a described, measured phenomenon with a known direction.
The system prompt text I actually use
Instead of describing it, here is the text itself, ready to copy and adapt to your dialect:
Write in spoken Damascene Levantine, the way a person writes on WhatsApp.
Strictly forbidden: sawfa, sa-, lam, lan, alladhi, allati, al-khass bik, yumkinuka an.
Use instead: rah, ma, yalli, tabaʿak, fik.
Short consecutive sentences, never one long compound sentence.
Keep any technical term exactly as the user wrote it in English.
Politeness does not mean MSA: be polite in the dialect.
Examples of the required tone:
“ahlein, akid ba’dar saʿdak. bas khabirni shu sar bil-zabt?”
“tamam, rah ba’atlak al-malaff khilal saʿa. iza fi shi na’is ihkili.”
“wallah ma intabaht la-hal-nu’ta, maʿak ha’. rah ʿaddilha halla’.”
The decisive element is the ban list. Forbidding four forms by name is stronger than ten sentences describing the dialect, because a model treats a ban as a constraint and a description as a suggestion. The second element is the three examples: a model imitates a tone it can see far better than it executes a description it reads.
How to measure each element’s effect
Run the same conversation four times, changing the prompt setup each time, and log one number only: the turn where the first structural leak appeared, meaning the first sawfa, lam, or fully MSA construction.
The ordering I expect from experience, and one worth verifying yourself because it varies by model:
| Setup | What I typically observe |
|---|---|
| Plain request: “reply in Syrian” | Early leakage, often by turn two or three |
| Request plus a verbal description of the dialect | Slight improvement, description alone does not hold |
| Request plus three written examples | Noticeable improvement, imitation beats description |
| Examples plus an explicit ban list | Clearly best, leakage pushed to later turns |
The practical point: you can improve this number yourself without waiting for a new release. It is one of the few places in model evaluation where the decision is in your hands rather than the lab’s.
From the test notebook: the deletion test applied to two texts
The deletion test is my fastest tool, and here it is applied in full to two texts.
Text one, output from a model asked to reply in Levantine:
“ahlan fik, shu ba’dar saʿdak? halla’ sawfa aqum bi-murajaʿat talabak, wa lam atamakkan min dhalik sabiqan li-anna al-nizam al-khass bina kana mutawaqqifan.”
Delete the explicitly dialectal words: ahlan fik, shu, halla’. What remains?
“ba’dar saʿdak? sawfa aqum bi-murajaʿat talabak, wa lam atamakkan min dhalik sabiqan li-anna al-nizam al-khass bina kana mutawaqqifan.”
Essentially clean MSA. Verdict: full leakage.
Text two, written by a native speaker:
“ahlein, kif ba’dar saʿdak? halla’ bshuf talabak, ma ‘dirt shufo abl li-ann al-nizam tabaʿna kan wa’iʿ.”
Delete the same words: ahlein, halla’. What remains?
“kif ba’dar saʿdak? bshuf talabak, ma ‘dirt shufo abl li-ann al-nizam tabaʿna kan wa’iʿ.”
Text that is broken as MSA, because the dialect sits inside its structure: bshuf not sa-ara, ma ‘dirt not lam atamakkan, tabaʿna not al-khass bina, wa’iʿ not mutawaqqif. Verdict: authentic dialect.
This test takes ten seconds and needs no tooling, and it is more accurate than any lexical metric I have seen.
The five indicators I count in every dialectal text
| Indicator | Dialectal | Leaked MSA |
|---|---|---|
| Future | rah, ha- | sawfa, sa- |
| Negation | ma | lam, la, lan |
| Relative pronoun | yalli | alladhi, allati |
| Possession | tabaʿi, ili | al-khass bi, ladayya |
| Present verb | bshuf, bruh | ara, adhhab |
Five indicators, a quick count. If three appear in their MSA form inside a text that is supposed to be dialectal, the model did not write a dialect.
Why examples beat descriptions
I have tried both repeatedly and the difference is clear: a model imitates what it sees better than it executes what it reads. Describing the dialect with “write in spoken Levantine” produces a far weaker result than three short examples written by a native speaker inside the system prompt.
The reason is structural: a description is processed as an instruction open to interpretation, an example is processed as a pattern to follow. So I always put the examples before the rules, never after.
Sources
- HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
- Open Arabic LLM Leaderboard v2: huggingface.co/blog/leaderboard-arabic-v2
- Llama 4 details and variants: en.wikipedia.org/wiki/Llama_(language_model)
- ALLaM 34B evaluation and the MSA reversion pattern: arxiv.org/abs/2508.17378
- AL-QASIDA, the correlation between dialect failure and MSA reversion: arxiv.org/abs/2412.04193