Fourteen Thousand Questions From Our Own Schools: The Exam That Exposed the Translated Numbers

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Paper released: February 20, 2024 | Published: March 5, 2024

The numbers first

Metric Result What it measures Source
Number of questions 14,575 multiple choice The size of the exam ArabicMMLU paper, Findings of ACL 2024
Number of tasks 40 Breadth of domains Same paper
Source countries 8 Arab countries Geographic spread Same paper
The countries Morocco, Egypt, Jordan, Palestine, Lebanon, UAE, Kuwait, Saudi Arabia Where the questions came from Same paper
Share of machine translation Zero What separates it from everything before Same paper
Models evaluated 35 The scale of the survey Same paper
GPT-4 72.5% The highest of all Same paper
GPT-3.5 57.7% The generational gap: 14.8 points Same paper
Jais-chat 30B 62.3% The best open model Same paper
Jais-chat 13B 54.8% The smaller version Same paper
LLaMA2-13B 36.1% A global model with no declared Arabic Same paper
Falcon 40B 34.8% An Arab-built model Same paper
Random guessing 29.0% The floor Same paper
Education level distribution Primary 22.2%, middle 12.2%, high 34%, university 6.1%, the rest unclassified The structure of the exam Same paper

The number I want you to carry away: 72.5% for GPT-4 here, against 80% on translated MMLU. The native Arabic exam is seven and a half points harder for the same model.


The results map, from floor to ceiling

ArabicMMLU, average accuracy
Random guessing
29.0%
Jais-13B base
32.2%
Falcon 40B
34.8%
LLaMA2-13B
36.1%
Jais-chat 13B
54.8%
GPT-3.5
57.7%
Jais-chat 30B
62.3%
GPT-4
72.5%

Note the three rows immediately above guessing. Three extremely well known models stand between three and seven points from a dice roll.


Benchmark card

Item Detail
Name ArabicMMLU
Team Koto, Li, Shatnawi, Doughman, Sadallah, Alraeesi, Almubarak, Alyafeai, Sengupta, Shehata, Habash, Nakov and Baldwin
Paper date February 2024
Venue Findings of ACL 2024, pages 5622 to 5640
Language Modern Standard Arabic
Source Real school examinations
Method Constructed in collaboration with native speakers in the region
What it does not cover Dialects, speech, pragmatics, free generation

The ARAB-LENS reading: three things this exam exposed

First: translation was flattering the picture.

This is the most important conclusion. The same model, GPT-4, scores 80% on a machine-translated exam and 72.5% on one written in Arabic to begin with. The difference is seven and a half points, and the paper itself references the translated number and compares against it.

Seven and a half points is not a detail. In a model ranking, seven points can mean three places. And if every Arab decision in 2023 was made on the basis of translated numbers, it was made on a prettier picture than reality.

Second: the gap between Arabic models and global ones is narrower than assumed, but in one specific direction.

Look: Jais-chat 30B scored 62.3% and GPT-3.5 scored 57.7%. An open Arabic model beat GPT-3.5 on a native Arabic exam. That is a big piece of news and it never got the coverage it deserved.

But look also: GPT-4 scored 72.5%, ten points above the best Arabic model. And that gap is not an Arabic gap. It is a general capability gap. GPT-4 is smarter, and intelligence transfers into Arabic.

Which gives us the rule that keeps recurring in this series: Arabic specialisation beats mid-size scale and does not beat large scale.

Third, and this is my criticism of the exam itself: ArabicMMLU is a major achievement, and I use it and trust it. But we must know its limits.

It is a multiple choice exam, in Modern Standard Arabic, drawn from school curricula. Which means three things it does not measure: it does not measure dialects, it does not measure free generation, and it does not measure pragmatics. A model scoring 72.5% here may be entirely unable to write a natural Levantine message, or to understand that khalleena nshoof means no.

And notice a detail in the table: the education level percentages sum to only 74.5%, with the remaining quarter unclassified. It is a small observation, but it reminds you that every benchmark, however good, has undefined regions. Anyone who quotes a number from a benchmark without reading its boundaries is quoting half the truth.


From the test notebook: building your own Arabic exam

ArabicMMLU measures school knowledge. Your product does not sell school knowledge. This is how I build an exam that measures what actually matters to you.

Step one: five categories, ten questions each.

Category An e-commerce example
Product knowledge “What is the difference between your warranty and your exchange policy?”
Dialect comprehension “The size came out small, I want to exchange it. How?”
Pragmatics “My order is a week late and nobody has replied to me.”
Cultural sensitivity “I want a gift for a bride, what do you recommend?”
Correct refusal “Give me the manager’s personal mobile number.”

That last category is the one everyone forgets. A good model refuses politely and in correct Arabic. It does not refuse in English and it does not comply.

Step two: write your questions in three dialects.

The same question in Levantine, Egyptian and Gulf. That triples the size of your exam and gives you the single most valuable fact you own: which dialect does your product actually serve?

Levantine Egyptian Gulf
“biddi arajjeʿ al-talabiyye” “ʿaayez arrajjaʿ al-order” “abi arajjeʿ al-talab”
“shu raqm al-shahne?” “eih raqm al-shahn?” “wesh raqm al-shahna?”
“ma wasalni shi la-hallaʾ” “magaleesh haaga lehadd delwaʾti” “ma wasalni shay lein al-heen”

Step three: grade on a three-point scale, not right or wrong.

Score Meaning
2 Correct and in the asker’s dialect
1 Correct but in formal Arabic or another dialect
0 Wrong, or in English, or an inappropriate refusal

The difference between 2 and 1 is the difference that determines customer satisfaction, and no global benchmark measures it.

Step four: keep it and do not change it.

Fifty fixed questions you run against every new model and every upgrade. After a year you have a real curve for your own product. And that, in my view, is a hundred times more useful than following global leaderboards.

Final item: the ruler question.

One question I put in every exam I build, because it separates quickly:

“A customer sent you a message saying: in shaa Allah I will drop by tomorrow. Do you book them an appointment or not?”

The correct answer: book them provisionally and confirm by message, because the phrase is a sincere intention without a firm commitment. A model that says “yes, book it” confidently or “no, they did not confirm” confidently has got it wrong both times. The correct answer contains a hedge.

Step five: add an “I do not know” category.

An addition I recommend strongly and nobody makes. Make it an acceptable answer for the model to say “I do not know” or “I need more information”.

Behaviour Score
Correct answer 2
Admitting ignorance when it does not know 1
Confidently wrong 0
Wrong but hedged 0.5

Why? Because a confident error costs more than an admission of ignorance in any real application. And global benchmarks never measure this, because multiple choice allows no “I do not know” box.


What distinguishes a good Arabic benchmark

After reading a number of Arabic benchmarks, these are the quality criteria I arrived at:

Property Why it matters Present in ArabicMMLU
Questions authored in Arabic No translation distortion Yes
Geographic diversity No single country dominates Yes, eight countries
Native human review Catches the oddities Yes
Dialect coverage Measures reality No
Open-ended rather than multiple choice Measures generation No
Periodic refresh Resists saturation Not declared
Full publication of the method Allows criticism Yes

Four out of seven. That is the best we have today, and it simultaneously shows how far there is still to go.

And the point I want to fix in place: rows four and five are the most important for any commercial product, and they are exactly what no Arabic benchmark measures today. So anyone building an Arabic product has no external reference for the two most important dimensions of their product, and no option but to build their own.


The practical verdict

Stop using translated MMLU as an Arabic reference. A native alternative has existed since February 2024, and it is available, free, and published at a peer-reviewed venue.

And know that 72.5% is a ceiling, not a floor. The best model in the world got more than a quarter of Arab school curriculum questions wrong. If your product depends on Arabic factual accuracy, build a verification layer.

Most important: a general benchmark tells you who is best overall. Your own exam tells you who is best for you. And the second one is what pays your bills.


Next in the series

Two weeks later Anthropic released the Claude 3 family, and its model card contains multilingual results. The next article goes looking for Arabic in that card, and finds it in exactly one place, which is not the place you would expect.


Sources

  • Koto et al., ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic, arXiv:2402.12840, February 2024
  • Peer-reviewed version: Findings of ACL 2024, pages 5622 to 5640: aclanthology.org/2024.findings-acl.334
  • OpenAI, GPT-4 Technical Report, arXiv:2303.08774, for the translated comparison