The Leaderboard That Corrected Itself in Public: How Arabic Models Get Graded

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Leaderboard launched: May 14, 2024 | Published: June 4, 2024

The numbers first

Metric Result What it measures Source
Leaderboard launch date May 14, 2024 Hugging Face with the Technology Innovation Institute Launch blog
The original AlGhafa benchmark Roughly a dozen datasets, mostly native Arabic The Arabic foundation Almazrouei et al., ArabicNLP 2023
The later extension Eleven additional translated datasets What came in afterwards Launch blog
ACVA and AceGPT benchmarks 58 datasets Culture and values Launch blog
Added benchmarks Translated MMLU and EXAMS Knowledge Launch blog
Human reviewer acceptance of AlGhafa’s machine-translated questions 58% only The quality of machine translation, per the benchmark’s own builders AlGhafa paper
Version two admission, first “Many tasks were direct translations from English, which frequently introduced linguistic and contextual mismatches” Published self-criticism OALL v2 blog, February 10, 2025
Second admission “A large portion of datasets and tasks originated from non-Arabic-speaking contexts” Published self-criticism Same source
Third admission “Some benchmarks … became less effective over time as models achieved near-perfect scores” Benchmark saturation Same source
The disclosed software bug Answer verification used the text of the options rather than their indices, disproportionately harming smaller models A grading fault that affected the published ranking Same source

That last row is the rarest thing in this field: a measuring body announcing that its own arithmetic was wrong.


What changed between the versions

Composition of the Arabic evaluation leaderboard
native Arabic
machine-translated
native Arabic
human-translated
machine-translated
removed
Version one, May 2024
Version two, February 2025

An explicit note: this chart illustrates the direction of change as the maintainers described it, not published percentages. They described the scale in words like “many” and “a large portion” and never published an exact proportion, and I will not invent one.


Leaderboard card

Item Detail
Name Open Arabic LLM Leaderboard
Institutions Hugging Face and the Technology Innovation Institute
Version one May 14, 2024
Version two February 10, 2025
Removed in version two Saturated and machine-translated tasks
Added Native Arabic MMLU with 40 tasks, human-translated MMLU with 57 tasks, MedinaQA, AraTrust, and ALRAGE for retrieval

The ARAB-LENS reading: why I respect this leaderboard more after its mistakes

I will state an opinion that may sound odd: version two of this leaderboard is more important than version one, and the most important thing in it is the admission.

Let us unpack what happened.

In May 2024 the first serious Arabic leaderboard launched. That in itself is an achievement: before it, anyone wanting to know which model was better in Arabic relied on impressions. The leaderboard gave the region a shared reference.

Then nine months later, its maintainers published an assessment of their own work and acknowledged three things.

First: many tasks were direct translations. Which is what I have repeated in every article of this series, and now the benchmark’s own builders say it about their benchmark. Not an accusation from outside but a diagnosis from within.

Second: a large portion of the content originated in non-Arab contexts, and “often failed to reflect real-world use cases or meet the practical needs of Arabic-speaking communities”. Read that sentence again. This is a leaderboard saying about itself that it was measuring something else.

Third, and the most striking: a software bug in the grading. The answer verification mechanism in the AlGhafa tasks used the option text itself rather than its numeric index. And the consequence, by their own description, was that it disproportionately harmed smaller models.

What does that mean in practice? It means a ranking published for nine months, on which teams built decisions and companies built plans, was affected by a fault in the code. And a small model may have looked weaker than it was.

Why do I respect this rather than attack it?

Because the alternative is far worse. The alternative is that the bug is found from outside, or never found at all, or silently patched in an update nobody reads. What happened here is that a team published its own critique in detail and rebuilt the instrument.

I write this entire series on one principle: a number without a method is not a number. And this leaderboard, in version two, did exactly what I argue for. It published its method, criticised it, and fixed it.

But two criticisms remain.

The first: nine months is a long time. Real decisions were taken during it on the basis of faulty numbers. And there is no mechanism to notify anyone who built on the old figures.

The second, and more general: the saturation problem. When models approach a perfect score on a benchmark, the benchmark stops distinguishing between them. Which means every Arabic benchmark has a shelf life, and building benchmarks is not a project you complete once but an ongoing commitment. The region still treats benchmarks as projects rather than as infrastructure.


From the test notebook: how to read any Arabic leaderboard

Seven questions I put to any leaderboard before I believe a number from it.

The question Why it matters The warning sign
What share of tasks is machine-translated? Translation measures something else If the answer is not published
When was the evaluation run? Models change under the same name Entries with no date
Which metric, acc or acc_norm? The gap can exceed twenty points One unlabelled column
Is this the official model or a community version? A fine-tune is not the model An unfamiliar owner name
Is the benchmark saturated? If everyone is above 90 it no longer discriminates Scores bunched at the top
Does it cover dialects? Most leaderboards are formal Arabic only No mention of dialect
Is the standard error published? A two point gap may be noise Numbers with no margins

Bonus item: the sensitivity check.

If you have access to two models a point or two apart in the ranking, run your own fifty question exam on both. In my experience, the gap you see on the leaderboard does not predict the gap you see on your data. And more than once the ordering flips entirely.

That is not a defect in the leaderboard. It is the nature of measurement: the leaderboard measures a general average, and you need performance on your specific case.

Final item: what no Arabic leaderboard measures even today.

Dimension Is it measured?
Formal Arabic, multiple choice Yes, well
General knowledge Yes
Written dialects Partly, and weakly
Spoken dialects No, separate leaderboards
Pragmatics and context No
Correct refusal in Arabic No
Consistency across sessions No
Long-form generation quality Rarely

Half that table is empty in terms of coverage. That is not a criticism of the maintainers but a description of the state of the field. Anyone building an Arabic product today is measuring half of what they need.

A second bonus item: how to build your own counter-benchmark.

General benchmarks saturate and age. Yours does not. Here is how to build a “counter-benchmark” in a day:

Step Detail
1 Collect twenty questions your current model got wrong
2 Classify them: grammar, dialect, knowledge, pragmatics, refusal
3 Write the reference answer for each yourself
4 Run them against every new model
5 Add a new question every month from customer complaints

Why specifically the questions it got wrong? Because a useful benchmark is one that discriminates, and questions everyone passes discriminate nothing. Which is precisely the saturation lesson the leaderboard admitted.


A timeline of Arabic evaluation

Date The event The effect
2020 to 2022 ARLUE and ARGEN The first native Arabic benchmarks
December 2023 AlGhafa A multiple choice benchmark, mostly native
February 2024 ArabicMMLU The first broad Arab school exam
May 2024 OALL version one The first shared reference
February 2025 OALL version two Self-criticism and removal of translated tasks

Five milestones in five years. That is faster than many assume and slower than we need.

And the observation I will close on: four of those five milestones were built by Arab teams or with Arab funding. Benchmarks are the one area in this entire series that the region genuinely leads. And I think that is the right place to invest: whoever owns the benchmark sets the direction of improvement.


The practical verdict

Use leaderboards for the first pass, not the final decision. A leaderboard shortens a list of thirty models to five. Choosing one of the five needs your own exam.

And store the date alongside the number. “Scored X on version one of the leaderboard in May 2024” is a correct sentence today. Without the date it becomes a wrong sentence after a single update.

Most important: when you see an organisation publish its own critique, trust it more rather than less. Anyone who never publishes mistakes either does not make them, which is impossible, or does not look for them.


Next in the series

Two days after this leaderboard, Cohere For AI released Aya 23, an open model covering twenty three languages. And Arabic is the first language on its list. The next article is about an experiment that chose depth over breadth, and what it means for your language to be one of twenty three rather than one of a hundred and one.


Sources

  • Hugging Face and TII, Introducing the Open Arabic LLM Leaderboard, May 14, 2024: huggingface.co/blog/leaderboard-arabic
  • Hugging Face, The Open Arabic LLM Leaderboard 2, February 10, 2025: huggingface.co/blog/leaderboard-arabic-v2
  • Almazrouei et al., AlGhafa Evaluation Benchmark for Arabic Language Models, ArabicNLP 2023, pages 244 to 275: aclanthology.org/2023.arabicnlp-1.21