The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Number of questions | 14,575 multiple choice | The size of the exam | ArabicMMLU paper, Findings of ACL 2024 |
| Number of tasks | 40 | Breadth of domains | Same paper |
| Source countries | 8 Arab countries | Geographic spread | Same paper |
| The countries | Morocco, Egypt, Jordan, Palestine, Lebanon, UAE, Kuwait, Saudi Arabia | Where the questions came from | Same paper |
| Share of machine translation | Zero | What separates it from everything before | Same paper |
| Models evaluated | 35 | The scale of the survey | Same paper |
| GPT-4 | 72.5% | The highest of all | Same paper |
| GPT-3.5 | 57.7% | The generational gap: 14.8 points | Same paper |
| Jais-chat 30B | 62.3% | The best open model | Same paper |
| Jais-chat 13B | 54.8% | The smaller version | Same paper |
| LLaMA2-13B | 36.1% | A global model with no declared Arabic | Same paper |
| Falcon 40B | 34.8% | An Arab-built model | Same paper |
| Random guessing | 29.0% | The floor | Same paper |
| Education level distribution | Primary 22.2%, middle 12.2%, high 34%, university 6.1%, the rest unclassified | The structure of the exam | Same paper |
The number I want you to carry away: 72.5% for GPT-4 here, against 80% on translated MMLU. The native Arabic exam is seven and a half points harder for the same model.
The results map, from floor to ceiling
Note the three rows immediately above guessing. Three extremely well known models stand between three and seven points from a dice roll.
Benchmark card
| Item | Detail |
|---|---|
| Name | ArabicMMLU |
| Team | Koto, Li, Shatnawi, Doughman, Sadallah, Alraeesi, Almubarak, Alyafeai, Sengupta, Shehata, Habash, Nakov and Baldwin |
| Paper date | February 2024 |
| Venue | Findings of ACL 2024, pages 5622 to 5640 |
| Language | Modern Standard Arabic |
| Source | Real school examinations |
| Method | Constructed in collaboration with native speakers in the region |
| What it does not cover | Dialects, speech, pragmatics, free generation |
The ARAB-LENS reading: three things this exam exposed
First: translation was flattering the picture.
This is the most important conclusion. The same model, GPT-4, scores 80% on a machine-translated exam and 72.5% on one written in Arabic to begin with. The difference is seven and a half points, and the paper itself references the translated number and compares against it.
Seven and a half points is not a detail. In a model ranking, seven points can mean three places. And if every Arab decision in 2023 was made on the basis of translated numbers, it was made on a prettier picture than reality.
Second: the gap between Arabic models and global ones is narrower than assumed, but in one specific direction.
Look: Jais-chat 30B scored 62.3% and GPT-3.5 scored 57.7%. An open Arabic model beat GPT-3.5 on a native Arabic exam. That is a big piece of news and it never got the coverage it deserved.
But look also: GPT-4 scored 72.5%, ten points above the best Arabic model. And that gap is not an Arabic gap. It is a general capability gap. GPT-4 is smarter, and intelligence transfers into Arabic.
Which gives us the rule that keeps recurring in this series: Arabic specialisation beats mid-size scale and does not beat large scale.
Third, and this is my criticism of the exam itself: ArabicMMLU is a major achievement, and I use it and trust it. But we must know its limits.
It is a multiple choice exam, in Modern Standard Arabic, drawn from school curricula. Which means three things it does not measure: it does not measure dialects, it does not measure free generation, and it does not measure pragmatics. A model scoring 72.5% here may be entirely unable to write a natural Levantine message, or to understand that khalleena nshoof means no.
And notice a detail in the table: the education level percentages sum to only 74.5%, with the remaining quarter unclassified. It is a small observation, but it reminds you that every benchmark, however good, has undefined regions. Anyone who quotes a number from a benchmark without reading its boundaries is quoting half the truth.
From the test notebook: building your own Arabic exam
ArabicMMLU measures school knowledge. Your product does not sell school knowledge. This is how I build an exam that measures what actually matters to you.
Step one: five categories, ten questions each.
| Category | An e-commerce example |
|---|---|
| Product knowledge | “What is the difference between your warranty and your exchange policy?” |
| Dialect comprehension | “The size came out small, I want to exchange it. How?” |
| Pragmatics | “My order is a week late and nobody has replied to me.” |
| Cultural sensitivity | “I want a gift for a bride, what do you recommend?” |
| Correct refusal | “Give me the manager’s personal mobile number.” |
That last category is the one everyone forgets. A good model refuses politely and in correct Arabic. It does not refuse in English and it does not comply.
Step two: write your questions in three dialects.
The same question in Levantine, Egyptian and Gulf. That triples the size of your exam and gives you the single most valuable fact you own: which dialect does your product actually serve?
| Levantine | Egyptian | Gulf |
|---|---|---|
| “biddi arajjeʿ al-talabiyye” | “ʿaayez arrajjaʿ al-order” | “abi arajjeʿ al-talab” |
| “shu raqm al-shahne?” | “eih raqm al-shahn?” | “wesh raqm al-shahna?” |
| “ma wasalni shi la-hallaʾ” | “magaleesh haaga lehadd delwaʾti” | “ma wasalni shay lein al-heen” |
Step three: grade on a three-point scale, not right or wrong.
| Score | Meaning |
|---|---|
| 2 | Correct and in the asker’s dialect |
| 1 | Correct but in formal Arabic or another dialect |
| 0 | Wrong, or in English, or an inappropriate refusal |
The difference between 2 and 1 is the difference that determines customer satisfaction, and no global benchmark measures it.
Step four: keep it and do not change it.
Fifty fixed questions you run against every new model and every upgrade. After a year you have a real curve for your own product. And that, in my view, is a hundred times more useful than following global leaderboards.
Final item: the ruler question.
One question I put in every exam I build, because it separates quickly:
“A customer sent you a message saying: in shaa Allah I will drop by tomorrow. Do you book them an appointment or not?”
The correct answer: book them provisionally and confirm by message, because the phrase is a sincere intention without a firm commitment. A model that says “yes, book it” confidently or “no, they did not confirm” confidently has got it wrong both times. The correct answer contains a hedge.
Step five: add an “I do not know” category.
An addition I recommend strongly and nobody makes. Make it an acceptable answer for the model to say “I do not know” or “I need more information”.
| Behaviour | Score |
|---|---|
| Correct answer | 2 |
| Admitting ignorance when it does not know | 1 |
| Confidently wrong | 0 |
| Wrong but hedged | 0.5 |
Why? Because a confident error costs more than an admission of ignorance in any real application. And global benchmarks never measure this, because multiple choice allows no “I do not know” box.
What distinguishes a good Arabic benchmark
After reading a number of Arabic benchmarks, these are the quality criteria I arrived at:
| Property | Why it matters | Present in ArabicMMLU |
|---|---|---|
| Questions authored in Arabic | No translation distortion | Yes |
| Geographic diversity | No single country dominates | Yes, eight countries |
| Native human review | Catches the oddities | Yes |
| Dialect coverage | Measures reality | No |
| Open-ended rather than multiple choice | Measures generation | No |
| Periodic refresh | Resists saturation | Not declared |
| Full publication of the method | Allows criticism | Yes |
Four out of seven. That is the best we have today, and it simultaneously shows how far there is still to go.
And the point I want to fix in place: rows four and five are the most important for any commercial product, and they are exactly what no Arabic benchmark measures today. So anyone building an Arabic product has no external reference for the two most important dimensions of their product, and no option but to build their own.
The practical verdict
Stop using translated MMLU as an Arabic reference. A native alternative has existed since February 2024, and it is available, free, and published at a peer-reviewed venue.
And know that 72.5% is a ceiling, not a floor. The best model in the world got more than a quarter of Arab school curriculum questions wrong. If your product depends on Arabic factual accuracy, build a verification layer.
Most important: a general benchmark tells you who is best overall. Your own exam tells you who is best for you. And the second one is what pays your bills.
Next in the series
Two weeks later Anthropic released the Claude 3 family, and its model card contains multilingual results. The next article goes looking for Arabic in that card, and finds it in exactly one place, which is not the place you would expect.
Sources
- Koto et al., ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic, arXiv:2402.12840, February 2024
- Peer-reviewed version: Findings of ACL 2024, pages 5622 to 5640: aclanthology.org/2024.findings-acl.334
- OpenAI, GPT-4 Technical Report, arXiv:2303.08774, for the translated comparison