Two Years of Arabic: What Actually Changed Since Claude 3.5 Sonnet and Gemini 2.5 Pro

Reading Time: 9 min
18
Baseline releases: October 2024 and March 25, 2025 | Published: April 8, 2025 Ehab Saleh | techkahwa.net

The numbers first

What changedBeforeAfterSource
OALL leaderboardVersion one, partly built on machine-translated benchmarksVersion two dropped the translations for native benchmarks: ArabicMMLU, MadinahQA, AraTrust, ALRAGEOALL v2
Arabic speech measurementNo multidialectal shared taskNADI 2025, the first multidialectal Arabic speech processing shared taskArabicNLP 2025
Culture measurementMerged with religious knowledge, or absentPalmX 2025 separates Arab culture from Islamic cultureArabicNLP 2025
Dialogue measurementNo large Arabic benchmarkShawarma Chats: 30,000 six-turn conversationsArabicNLP 2025
Dialect measurementScatteredDialectalArabicMMLU: 21,945 questions across five dialects and 32 domainsLREC 2026
Independent holistic leaderboardNoneHELM Arabic from Stanford with Arabic.AI, seven benchmarksStanford CRFM, December 2025
Arabic modelsJais, ALLaM 13B, AceGPTJais 2 at 70B, Fanar 2.0 at 27B multimodal, Falcon-H1-Arabic up to 34B, ALLaM 34BVendor announcements, 2025 and 2026
Growth of Arabic evaluation infrastructure
2024
one board, partly translated benchmarks
2025
native benchmarks, speech, culture, dialogue
2026
independent board + dialects + new Arabic models

That is real and rapid progress, and it should be said plainly: the infrastructure for measuring Arabic in 2026 is far better than it was in 2024. The credit belongs to Arab academic and industry teams working on budgets a fraction of those spent measuring English.

Now the other side.


What did not change

ProblemStatus todayEvidence
The dialect, MSA, English gapStill there47.7%, 51.9%, 62.8%, DialectalArabicMMLU
MSA is not a bridge to dialectConfirmed numericallyTranslating to MSA gains 0.3 points against 5.6 for English
Dialect generationUnmeasured on any public boardSurvey of OALL and HELM Arabic benchmarks
PragmaticsUnmeasuredSame survey
LevantineWeakest in available measurements46.6% in DialectalArabicMMLU, 2.73 of 5 in the independent ALLaM evaluation
Dialectal speech transcriptionMore than a third of words wrong35.68 WER in the best system, NADI 2025
Dialect stereotypingNo metric existsSurvey of public sources

That is the conclusion I reach after two years: what was measurable improved, and what is hard to measure stayed where it was. It is no accident that the hard-to-measure part is precisely what decides whether an Arabic product works for its user: dialect, tone, intent, and social custom.


The two baseline cards

FieldClaude 3.5 SonnetGemini 2.5 Pro
LabAnthropicGoogle
DateJune 2024, upgraded version October 22, 2024March 25, 2025, experimental
What they representedThe generation before extended reasoningThe beginning of the chain-of-thought generation
Their Arabic footprintGood MSA, weak dialectNoticeable improvement, dialect gap intact

One point worth stating about Gemini 2.5 Pro specifically: in the AHELM holistic evaluation of audio-language models, it recorded the highest overall performance with a mean win rate of 0.803 across ten dimensions. And the same evaluation detected a statistically significant gender bias in some of its speech tasks. That is an accurate summary of this whole period: a genuine leap in capability, with unresolved fairness problems.


How to measure progress yourself: the fixed baseline set

The most useful lesson from the past two years is this: build a fixed test set that never changes, and run every new model against it.

Ten items is enough:

  1. A dialectal phrase in two opposed contexts (dabbir halak).
  2. A social phrase in four situations (Allah yaʿtik al-ʿafiya).
  3. Converting an MSA sentence into three dialects with structural, not lexical, differences.
  4. The register ladder: the same message to six different addressees.
  5. A hospitality scenario measuring cultural reasoning.
  6. An invented custom measuring fabrication.
  7. An undecidable question measuring admission of ignorance.
  8. A six-turn dialogue measuring relationship reading.
  9. A paragraph with five grammatical errors measuring proofreading.
  10. An audio clip in your dialect measuring transcription and comprehension.

Store the results with their dates. After four models you will have something nobody else holds: a time series of model capability on your dialect, not on MSA, and not on the average of twenty two countries.

Why a fixed set beats a public leaderboard
the board measures average Arabic
→
you sell to one dialect
the board changes across versions
→
your set stays comparable
the board measures what is easy
→
you measure what matters

Three testable predictions

I write them down here to return to later, which in my view is the only kind of writing in this field worth publishing:

  1. Dialect generation will stay without a public benchmark for at least another year, because measuring it requires native speakers who cannot be replaced by a judging model.
  2. The MSA to dialect gap will narrow in specialized Arabic models before general ones, because data is the constraint and Arab institutions are better placed to collect it.
  3. Whoever first builds a Levantine gold set written by native speakers and withheld from training will own a reference everyone returns to. The scarcity here is not in models, it is in trustworthy data.

In the final article of the series

Speech, the weakest link in Arabic today: the full NADI 2025 numbers, a comparison between the transcribe-then-understand path and the integrated audio model, and what any team building an Arabic voice product should measure before launching it.


What was added to the map after this article

This is an article about change, so it is only fair to add what changed after it. Two important layers entered the map:

Layer one, measuring dialect fidelity. The AL-QASIDA study measured nine models across eight dialects on four dimensions: fidelity, understanding, quality and diglossia. It produced numbers that became reference points: most ADI2 scores below 50%, a 0.99 correlation between dialect failure and MSA production, and a human evaluation describing outputs as fluent and adequate and mostly not in the requested dialect.

Layer two, cleaning the benchmarks themselves. The QIMMA leaderboard in April 2026, with more than 52,000 samples across fourteen benchmarks and seven domains, added something that did not exist: a sample quality validation pipeline running before evaluation. Its revealing result was a 3.1% discard rate in ArabicMMLU.

The Arabic measurement map, updated
2024
▸ AL-QASIDA: first systematic dialect fidelity measurement
2025
▸ OALL v2: translated benchmarks removed
▸ NADI: first multidialectal speech task
▸ PalmX: culture separated from religion
▸ Shawarma Chats: first large dialogue benchmark
▸ HELM Arabic: first independent holistic board
2026
▸ DialectalArabicMMLU: five dialects, 32 domains
▸ QIMMA: sample validation and code evaluation

The predictions, and how to judge them

I wrote three predictions above. Here are the criteria for judging them, so they are not just talk:

PredictionConfirmed ifFalsified if
Dialect generation stays unmeasured another yearNo major board adds a human-judged dialect generation metricA public board adds one judged by native speakers
The gap narrows in Arabic models firstA new Arabic model clears 60% ADI2 before any general modelA global model gets there first
Whoever builds a Levantine gold set winsA Levantine reference set appears and researchers cite itLevantine remains without a reference set

I consider writing a prediction with a measurable criterion part of the professionalism: publishing an opinion about the future without saying when you would be wrong is publishing an impression, not an opinion.

What I am watching over the next twelve months

  1. Will Levantine enter the dialogue benchmarks? Shawarma Chats covered only MSA, Egyptian and Maghrebi.
  2. Will a dialect stereotyping metric appear? There is none today.
  3. Will Arabic reasoning chains be published as a dataset? The clearest current gap.
  4. Will the benchmarks stay clean? After QIMMA’s finding, sample quality became a legitimate question for every board.

The baseline set: the ten items I rerun on every new model

The most important decision I made over the past two years was to fix a test set that never changes. Here it is in full, and it is what I run against every new release on the day it ships:

#ItemWhat it measuresThe answer I accept
1dabbir halak in a friend and a manager situationContext sensitivityTwo entirely different meanings
2Allah yaʿtik al-ʿafiya in four situationsSocial functionAcknowledgment, closing, thanks, sarcasm
3Converting a sentence into three dialectsStructure, not vocabularyDifferent negation, future and syntax
4The same message to six addresseesRegisterDifferent length and directness
5The hospitality scenarioCultural reasoningUnderstanding insistence as generosity
6An invented custom in a real cityFabricationAdmitting ignorance
7A one-word WhatsApp messageAcknowledging ambiguityRefusing to decide
8A six-turn dialogueRelationship readingA close relationship, not manager and employee
9A paragraph with five grammatical errorsProofreadingFinding the errors with reasoning
10An audio recording in my dialectTranscription and comprehensionAnswering about meaning, not copying text

Ten items, half an hour of work, and a result comparable across years. No leaderboard gives you that, because leaderboards change their benchmarks between versions, which destroys temporal comparability.

What actually changed in these items over two years

I write this as a personal observation from running the same set repeatedly, not as a measured figure:

Clearly improved: item nine, grammatical proofreading. Models now find errors and explain them rather than silently rewriting the paragraph. And item two, the social function of common phrases, improved as well, though sarcasm remains the hardest part.

Not much improved: item three, structural conversion between dialects. The output is still different vocabulary on identical structure. And item six, the invented custom, still catches many strong models.

Unchanged: the second half of item ten. Transcription improved, and understanding intent from spoken language has barely moved.

These observations are exactly what makes a fixed set more valuable than a shifting leaderboard: you are not only measuring the model, you are measuring the direction of the field.

Sources

  • Open Arabic LLM Leaderboard v2 and the removal of translated benchmarks: huggingface.co/blog/leaderboard-arabic-v2
  • HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • NADI 2025: arxiv.org/abs/2509.02038
  • PalmX 2025: arxiv.org/abs/2509.02550
  • Shawarma Chats, ArabicNLP 2025: aclanthology.org/2025.arabicnlp-main.39
  • DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
  • AHELM, holistic evaluation of audio-language models: arxiv.org/pdf/2508.21376
  • AL-QASIDA: arxiv.org/abs/2412.04193
  • QIMMA Arabic leaderboard: huggingface.co/blog/tiiuae/qimma-arabic-leaderboard