4.74 in MSA and 2.73 in Levantine: The Number That Sums Up the Arabic Evaluation Crisis

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: August 2025 | Published: September 14, 2025

The numbers first

An independent evaluation of ALLaM 34B through the HUMAIN Chat interface, judged by three separate frontier models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4), over 23 prompts repeated five times each for 115 total responses:

CategoryScore out of 5Source
Code-switching and generation4.92arXiv 2508.17378, August 2025
Knowledge4.77Same source
Modern Standard Arabic4.74Same source
Reasoning4.64Same source
Safety4.54Same source
Dialect handling4.21Same source

And now the table everyone writing about Arabic AI should read:

DialectScore out of 5Source
Najdiabout 3.8arXiv 2508.17378
Hijaziabout 3.8Same source
Egyptianabout 3.8Same source
Moroccan3.3Same source
Levantine2.73Same source
The gap inside a single model
MSA
4.74
Najdi
3.80
Moroccan
3.30
Levantine
2.73
Two models inside one: excellent in MSA,
failing in Levantine

Two full points between MSA and Levantine, in the same model, in the same session, judged the same way. That number is the clearest published evidence for the premise ARAB-LENS is built on: Arabic is not one language and cannot be scored with one number.


Why Levantine specifically fails

The most important observation in that evaluation is not the score, it is the failure mode the researchers described: the model reverts to MSA formalism when facing a dialectal prompt, especially in varieties less represented in its data, and occasionally breaks character entirely, answering with something resembling an English search result.

That is precisely what I call MSA leakage, the failure pattern I give its own metric inside the ARAB-DIALECT axis. And the cause is structural, not random:

a model built with Saudi funding and Saudi data
        ↓
excellent representation of Najdi, Hijazi and formal MSA
        ↓
weaker representation of Levantine and Moroccan
        ↓
when the dialect is beyond it, it falls back on what it does well: MSA

There is no moral failing in this. Every model reflects its data, and data reflects who funded it and where it was collected. But the practical consequence should be stated plainly: a national model that is excellent inside its own context can be weak outside its dialectal borders, and anyone building for Damascus, Amman or Beirut cannot import the headline number as is.


Model card

FieldDetail
NameALLaM 34B
LabsHUMAIN and the Saudi Data and AI Authority
ReleasedAugust 2025, alongside the HUMAIN Chat app
PredecessorALLaM 13B, May 2024 on IBM watsonx, then Azure in September 2024
Built asA foundation model trained from scratch, with forty PhD researchers and proprietary Saudi data
Commonly cited forLeading Arabic models on ArabicMMLU
Axis under testKnowledge versus dialect

The paradox: ArabicMMLU is not an Arabic test

ALLaM is widely credited with topping ArabicMMLU, a respectable and genuinely native benchmark built from real Arabic school exam questions rather than translations from English, and that is an achievement worth stating.

But let us be precise about what it measures: ArabicMMLU measures knowledge expressed in Arabic, not command of Arabic. A biology question written in Arabic measures biology. A question on Islamic history written in Arabic measures history. The language here is the container, not the subject.

Which is how the same model can hold a leading position in Arabic knowledge and a 2.73 in Levantine with no contradiction at all. The two measurements are measuring entirely different things:

BenchmarkActually measuresDoes not measure
ArabicMMLUAcademic knowledge expressed in ArabicNatural language production
MadinahQAKnowledge of Arabic grammar rulesApplying those rules in generated text
AraTrustSafetySocial appropriateness
UI-level evaluation (arXiv 2508.17378)Dialect fidelity, fluency, instruction followingBreadth of knowledge

I consider that comparison the core of this article: there is no bad benchmark in that list, only bad use of benchmarks. When someone says “this is the best model in Arabic” on the strength of ArabicMMLU alone, they are transporting a correct number to a question nobody asked.


UI-level evaluation: why I like it and where I am careful

The evaluation this article rests on ran through the user interface, not an API. Meaning the researchers tested what a user actually meets: the model plus its safety layer plus its system prompt plus everything the company stacks on top.

That is a real advantage, because what reaches the user is not the raw model. It is also a limit worth naming: the result belongs to the product, not the model, and it can shift with a system prompt change alone.

And a third point that concerns me as an evaluator: models were used as judges of a model. That is useful for scale and risky on its own, because a judging model may reward beautiful Arabic rather than socially correct Arabic, which is exactly the fluent-and-wrong trap. So ARAB-LENS holds a fixed rule: the final verdict on dialect fidelity belongs to a native speaker, and the judging model is a supporting layer, never the source of truth. To their credit, the researchers here added human validation on top of the automated scoring.


The test: does your model know its dialectal limits?

Item 1, an explicit dialect request. Ask for a reply in Levantine, then run the deletion test: remove every explicitly dialectal word and read what remains. If it is perfectly sound MSA, leakage has occurred.

Item 2, breaking character. Extend the conversation five turns in dialect and watch which turn it reverts on. I treat the reversion turn as a metric in itself, and call it dialectal stamina.

Item 3, internal comparison. Ask for the same reply in Najdi and then in Levantine, and compare naturalness. The gap you see is the gap between 3.8 and 2.73, except measured with your own ear rather than a judging model’s.

Item 4, admitting the limit. Ask it directly: which dialects do you handle well, and which are you weak in? A model that answers honestly is more useful than one claiming mastery of twenty two dialects.


The practical verdict

  1. Choose your model by your user’s dialect, not by the leaderboard. A model that leads in Najdi can be the worst option for a Levantine product, and the reverse.
  2. Do not use ArabicMMLU as evidence of Arabic writing quality. Use it as evidence of knowledge, which is what it measures.
  3. Measure dialectal stamina. The number of turns before a model reverts to MSA is an excellent practical indicator and trivial to compute.
  4. Never let a model be the sole judge of your dialect. Put a native speaker in the loop, even on just ten samples.

Next in the series

Back to general models, specifically Claude Opus 4.6, holder of one of the highest published Arabic scores on Global-MMLU-Lite, and an axis almost nobody measures: multi-turn conversation, and what happens to Arabic after the fifth turn.


Is this one model’s problem? No

The 2.73 could be read as a weakness in one particular model. Research honesty requires placing it in its wider context, and that context exists in the AL-QASIDA study, which measured nine models across eight dialects including Syrian and Palestinian.

Its results:

MeasureResult
ADI2 in monolingual settingsBelow 50% for nearly every model
GPT-4o cross-linguallyDid not clear 13% on half the dialects
ALDi and ADI2 correlation0.99
Human evaluationFluent, adequate responses, mostly not in the requested dialect

The 0.99 correlation is the key: when a model fails to produce the dialect, it produces MSA specifically, not another dialect and not a hybrid. Which means what the ALLaM evaluation caught in one model is general behavior across the whole category.

That changes the conclusion: the gap between 4.74 and 2.73 is not a defect in a Saudi model, it is the distance between what all of these models command and what none of them do. The Saudi model led in Najdi because its data is Najdi, and any other model will lead in whichever dialect dominates its data and fall back to MSA outside it.

Choosing your model by market rather than by leaderboard

Your marketWhat to ask about first
Saudi Arabia and the GulfNajdi and Hijazi representation in the data, where local models lead
The LevantTest it yourself, it is the weakest across every available measurement
EgyptThe best represented dialect in most benchmarks
The MaghrebThe hardest, with French code-switching on top
A cross-country audienceA strong MSA model with a dialect layer anchored by examples, not a model claiming all dialects

Measuring dialectal stamina in practice

TurnWhat to log
1Did it start in dialect at all?
2Is the structure still dialectal, or only the vocabulary?
3First appearance of an MSA form (sawfa, lam, alladhi)
4Did it return after your reminder, or continue in MSA?
5Final state

The number you leave with is the turn at which leakage occurred. In my experience a single reminder buys back two or three turns, then the model reverts again. Which is why the fix is not reminding, it is examples inside the system prompt.


From the test notebook: the equivalence table I use

When measuring a model’s command of a specific dialect, I do not ask it to “write in Najdi”. I give it one sentence and ask for it in five dialects, then compare against the equivalence table I keep:

MeaningNajdiLevantineEgyptianMoroccanEmirati Gulf
whatweshshuehashnushu or wesh
nowal-hinhalla’dilwa’tidabaal-hin
I wantabghabiddiʿayizbghitaba or abi
goodzeinmnihkwayyismezyanzein
how are youkeifak, shlonakkeifakizzayaklabasshhalak
I do not havema ʿindima ʿindimish ʿandima ʿandishma ʿindi

A model that hands me five sentences with different vocabulary and one shared structure did not translate a dialect. Dialects do not differ only in the lexicon: Egyptian and Moroccan negation wraps the verb in ma and sh, Levantine uses ma before it alone, and Gulf uses ma and mu in different positions.

Rural versus urban testing inside one dialect

This is what no leaderboard measures, and it is what matters most to a real product:

“gultillo ya khayy la tit’akhkhar, gal-li insha’Allah.”

This is clearly northern rural Jordanian: qaf as g, ya khayy a common northern vocative, and fast assimilation. A model trained on urban Amman treats it as “not Levantine”, and one trained on Gulf data sends it to Najd.

The correct answer: Jordanian Levantine with a northern Bedouin feature. That is the item I use to measure whether a model knows dialect is a spectrum rather than a box.

Why this matters when choosing a model

The published result says 3.8 for Najdi and 2.73 for Levantine in one model. That does not mean the model is weak, it means its data is Najdi. Measure a model trained on Levantine data and the ordering flips.

So the right question when selecting a model is not “which is better at Arabic” but “which dialect dominates its data, and is that my users’ dialect”. You can answer that in ten minutes with the table above, and no leaderboard in the world answers it for you.

Sources

  • UI-level evaluation of ALLaM 34B via HUMAIN Chat: arxiv.org/abs/2508.17378
  • HUMAIN announcement of the ALLaM 34B launch: middleeastainews.com
  • Open Arabic LLM Leaderboard v2 and its benchmarks: huggingface.co/blog/leaderboard-arabic-v2
  • HELM Arabic, Stanford CRFM: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • AL-QASIDA, dialect fidelity measurement: arxiv.org/abs/2412.04193