The numbers first
| What changed | Before | After | Source |
|---|---|---|---|
| OALL leaderboard | Version one, partly built on machine-translated benchmarks | Version two dropped the translations for native benchmarks: ArabicMMLU, MadinahQA, AraTrust, ALRAGE | OALL v2 |
| Arabic speech measurement | No multidialectal shared task | NADI 2025, the first multidialectal Arabic speech processing shared task | ArabicNLP 2025 |
| Culture measurement | Merged with religious knowledge, or absent | PalmX 2025 separates Arab culture from Islamic culture | ArabicNLP 2025 |
| Dialogue measurement | No large Arabic benchmark | Shawarma Chats: 30,000 six-turn conversations | ArabicNLP 2025 |
| Dialect measurement | Scattered | DialectalArabicMMLU: 21,945 questions across five dialects and 32 domains | LREC 2026 |
| Independent holistic leaderboard | None | HELM Arabic from Stanford with Arabic.AI, seven benchmarks | Stanford CRFM, December 2025 |
| Arabic models | Jais, ALLaM 13B, AceGPT | Jais 2 at 70B, Fanar 2.0 at 27B multimodal, Falcon-H1-Arabic up to 34B, ALLaM 34B | Vendor announcements, 2025 and 2026 |
That is real and rapid progress, and it should be said plainly: the infrastructure for measuring Arabic in 2026 is far better than it was in 2024. The credit belongs to Arab academic and industry teams working on budgets a fraction of those spent measuring English.
Now the other side.
What did not change
| Problem | Status today | Evidence |
|---|---|---|
| The dialect, MSA, English gap | Still there | 47.7%, 51.9%, 62.8%, DialectalArabicMMLU |
| MSA is not a bridge to dialect | Confirmed numerically | Translating to MSA gains 0.3 points against 5.6 for English |
| Dialect generation | Unmeasured on any public board | Survey of OALL and HELM Arabic benchmarks |
| Pragmatics | Unmeasured | Same survey |
| Levantine | Weakest in available measurements | 46.6% in DialectalArabicMMLU, 2.73 of 5 in the independent ALLaM evaluation |
| Dialectal speech transcription | More than a third of words wrong | 35.68 WER in the best system, NADI 2025 |
| Dialect stereotyping | No metric exists | Survey of public sources |
That is the conclusion I reach after two years: what was measurable improved, and what is hard to measure stayed where it was. It is no accident that the hard-to-measure part is precisely what decides whether an Arabic product works for its user: dialect, tone, intent, and social custom.
The two baseline cards
| Field | Claude 3.5 Sonnet | Gemini 2.5 Pro |
|---|---|---|
| Lab | Anthropic | |
| Date | June 2024, upgraded version October 22, 2024 | March 25, 2025, experimental |
| What they represented | The generation before extended reasoning | The beginning of the chain-of-thought generation |
| Their Arabic footprint | Good MSA, weak dialect | Noticeable improvement, dialect gap intact |
One point worth stating about Gemini 2.5 Pro specifically: in the AHELM holistic evaluation of audio-language models, it recorded the highest overall performance with a mean win rate of 0.803 across ten dimensions. And the same evaluation detected a statistically significant gender bias in some of its speech tasks. That is an accurate summary of this whole period: a genuine leap in capability, with unresolved fairness problems.
How to measure progress yourself: the fixed baseline set
The most useful lesson from the past two years is this: build a fixed test set that never changes, and run every new model against it.
Ten items is enough:
- A dialectal phrase in two opposed contexts (dabbir halak).
- A social phrase in four situations (Allah yaʿtik al-ʿafiya).
- Converting an MSA sentence into three dialects with structural, not lexical, differences.
- The register ladder: the same message to six different addressees.
- A hospitality scenario measuring cultural reasoning.
- An invented custom measuring fabrication.
- An undecidable question measuring admission of ignorance.
- A six-turn dialogue measuring relationship reading.
- A paragraph with five grammatical errors measuring proofreading.
- An audio clip in your dialect measuring transcription and comprehension.
Store the results with their dates. After four models you will have something nobody else holds: a time series of model capability on your dialect, not on MSA, and not on the average of twenty two countries.
Three testable predictions
I write them down here to return to later, which in my view is the only kind of writing in this field worth publishing:
- Dialect generation will stay without a public benchmark for at least another year, because measuring it requires native speakers who cannot be replaced by a judging model.
- The MSA to dialect gap will narrow in specialized Arabic models before general ones, because data is the constraint and Arab institutions are better placed to collect it.
- Whoever first builds a Levantine gold set written by native speakers and withheld from training will own a reference everyone returns to. The scarcity here is not in models, it is in trustworthy data.
In the final article of the series
Speech, the weakest link in Arabic today: the full NADI 2025 numbers, a comparison between the transcribe-then-understand path and the integrated audio model, and what any team building an Arabic voice product should measure before launching it.
What was added to the map after this article
This is an article about change, so it is only fair to add what changed after it. Two important layers entered the map:
Layer one, measuring dialect fidelity. The AL-QASIDA study measured nine models across eight dialects on four dimensions: fidelity, understanding, quality and diglossia. It produced numbers that became reference points: most ADI2 scores below 50%, a 0.99 correlation between dialect failure and MSA production, and a human evaluation describing outputs as fluent and adequate and mostly not in the requested dialect.
Layer two, cleaning the benchmarks themselves. The QIMMA leaderboard in April 2026, with more than 52,000 samples across fourteen benchmarks and seven domains, added something that did not exist: a sample quality validation pipeline running before evaluation. Its revealing result was a 3.1% discard rate in ArabicMMLU.
The predictions, and how to judge them
I wrote three predictions above. Here are the criteria for judging them, so they are not just talk:
| Prediction | Confirmed if | Falsified if |
|---|---|---|
| Dialect generation stays unmeasured another year | No major board adds a human-judged dialect generation metric | A public board adds one judged by native speakers |
| The gap narrows in Arabic models first | A new Arabic model clears 60% ADI2 before any general model | A global model gets there first |
| Whoever builds a Levantine gold set wins | A Levantine reference set appears and researchers cite it | Levantine remains without a reference set |
I consider writing a prediction with a measurable criterion part of the professionalism: publishing an opinion about the future without saying when you would be wrong is publishing an impression, not an opinion.
What I am watching over the next twelve months
- Will Levantine enter the dialogue benchmarks? Shawarma Chats covered only MSA, Egyptian and Maghrebi.
- Will a dialect stereotyping metric appear? There is none today.
- Will Arabic reasoning chains be published as a dataset? The clearest current gap.
- Will the benchmarks stay clean? After QIMMA’s finding, sample quality became a legitimate question for every board.
The baseline set: the ten items I rerun on every new model
The most important decision I made over the past two years was to fix a test set that never changes. Here it is in full, and it is what I run against every new release on the day it ships:
| # | Item | What it measures | The answer I accept |
|---|---|---|---|
| 1 | dabbir halak in a friend and a manager situation | Context sensitivity | Two entirely different meanings |
| 2 | Allah yaʿtik al-ʿafiya in four situations | Social function | Acknowledgment, closing, thanks, sarcasm |
| 3 | Converting a sentence into three dialects | Structure, not vocabulary | Different negation, future and syntax |
| 4 | The same message to six addressees | Register | Different length and directness |
| 5 | The hospitality scenario | Cultural reasoning | Understanding insistence as generosity |
| 6 | An invented custom in a real city | Fabrication | Admitting ignorance |
| 7 | A one-word WhatsApp message | Acknowledging ambiguity | Refusing to decide |
| 8 | A six-turn dialogue | Relationship reading | A close relationship, not manager and employee |
| 9 | A paragraph with five grammatical errors | Proofreading | Finding the errors with reasoning |
| 10 | An audio recording in my dialect | Transcription and comprehension | Answering about meaning, not copying text |
Ten items, half an hour of work, and a result comparable across years. No leaderboard gives you that, because leaderboards change their benchmarks between versions, which destroys temporal comparability.
What actually changed in these items over two years
I write this as a personal observation from running the same set repeatedly, not as a measured figure:
Clearly improved: item nine, grammatical proofreading. Models now find errors and explain them rather than silently rewriting the paragraph. And item two, the social function of common phrases, improved as well, though sarcasm remains the hardest part.
Not much improved: item three, structural conversion between dialects. The output is still different vocabulary on identical structure. And item six, the invented custom, still catches many strong models.
Unchanged: the second half of item ten. Transcription improved, and understanding intent from spoken language has barely moved.
These observations are exactly what makes a fixed set more valuable than a shifting leaderboard: you are not only measuring the model, you are measuring the direction of the field.
Sources
- Open Arabic LLM Leaderboard v2 and the removal of translated benchmarks: huggingface.co/blog/leaderboard-arabic-v2
- HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
- NADI 2025: arxiv.org/abs/2509.02038
- PalmX 2025: arxiv.org/abs/2509.02550
- Shawarma Chats, ArabicNLP 2025: aclanthology.org/2025.arabicnlp-main.39
- DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
- AHELM, holistic evaluation of audio-language models: arxiv.org/pdf/2508.21376
- AL-QASIDA: arxiv.org/abs/2412.04193
- QIMMA Arabic leaderboard: huggingface.co/blog/tiiuae/qimma-arabic-leaderboard