Eighty Percent in Arabic: The First Official Number, and the First Misleading One

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: March 14, 2023 | Published: March 28, 2023

The numbers first

Metric Result What it measures Source
GPT-4 on MMLU in Arabic 80% General knowledge, on a machine-translated exam GPT-4 Technical Report, Figure 5, March 2023
GPT-4 on MMLU in English, same condition 85.5% The reference the Arabic number is set against Same report, Figure 5
How the Arabic version was produced Machine translation via Azure Translate How the Arabic exam was built Same report
Examples given in the prompt 3, not 5 An easier measurement condition than the standard Same report
GPT-4 on English MMLU, standard condition 86.4% at five examples A completely different number, not to be mixed with the above Same report, Table 2
GPT-3.5 on English MMLU 70.0% The size of the generational jump Same report, Table 2
Languages where GPT-4 beat the prior English state of the art 24 of 26 The scale of the multilingual achievement Same report
GPT-4 on ArabicMMLU, a native Arabic exam 72.5% Nearly the same test, but with questions written in Arabic to begin with Koto et al., ArabicMMLU, February 2024

The first row and the last row are the article. Eighty on a translated exam, seventy two and a half on a real Arabic one. Seven and a half points disappeared the moment translation was replaced by the original.


Where the seven points went

GPT-4 on knowledge exams, by how the exam was built
English original, three-shot
85.5%
Arabic machine-translated
80.0%
Arabic native, ArabicMMLU
72.5%

The gap between the second bar and the third is not a gap in intelligence. It is a gap in what the exam measures.


Model card

Item Detail
Model GPT-4
Lab OpenAI
Release date March 14, 2023
Technical report arXiv:2303.08774
What was published about Arabic One number inside a chart, with no table and no breakdown
What was not published Anything about dialects, morphology, culture or diacritics
Tokenizer The same cl100k_base, so the letter tax is unchanged

The ARAB-LENS reading: three things hidden inside one number

First, and most serious: the exam is machine-translated.

The report says it plainly. The language versions of MMLU were produced with Azure Translate. Which means the Arabic question is not an Arabic question. It is an American question wearing Arabic letters.

Why does that matter? For three reasons that have accumulated in my notes from reading dozens of these benchmarks:

One, machine translation simplifies. It produces short sentences, direct structure and common vocabulary. So the Arabic exam is linguistically easier than its English counterpart.

Two, machine translation sometimes leaks the answer. When a correct option and a distractor are both translated, the correct one often comes out in cleaner phrasing because the source text was cleaner. Models pick that pattern up.

Three, and this is the big one, the content itself is American. Questions about American law, the American school system, European history. A model that answers them in Arabic has not shown that it knows Arabic. It has shown that it knows English and can type in Arabic.

Second: three examples, not five.

The standard MMLU condition is five worked examples. Figure 5 used three. That is not cheating, the report discloses it. But it means the Arabic number is not directly comparable to any other MMLU figure you read elsewhere. I have watched this confusion repeat in article after article: someone takes 80% from here and 86.4% from Table 2 and puts them in one sentence, when they are two different measurements under two different conditions.

Third: the number lives inside a picture.

The Arabic value does not exist as text in the report. It is the height of a bar in a chart. I cannot confirm whether it is 80.0 or 79.8 or 80.4. The one independent source that quotes it, the ArabicMMLU paper, writes it as a rounded “80%”. So when you read 80% anywhere, know that it has been rounded off a graph.

And here is my verdict: this is not an Arabic evaluation, it is a polite acknowledgement that Arabic exists. That is real progress compared to the total silence before it, but it is not what it appears to be.

The proof arrived eleven months later. When an academic team built a native Arabic exam from real school questions across eight Arab countries, the same GPT-4 scored 72.5%. Eighty becomes seventy two and a half the moment the questions stop being translated American ones.


From the test notebook: how to tell a benchmark was translated

These are the checks I run on any Arabic benchmark before I believe its number.

Item one: hunt for the American residue.

Open a sample of a hundred questions and count:

What you look for What its presence means
Names of American states or cities The exam was translated, not authored
GPA or SAT grading systems A non-Arab educational context
Imperial units, miles and pounds Literal translation with no localisation
Gregorian-only dates inside religious questions No Arab reviewer ever read it
“The president” with no qualifier, meaning the American one An unstated contextual assumption

Item two: the inverted sentence test.

Machine translation produces English word order in Arabic words. Its clearest signature is a nominal sentence where a verbal one belongs.

Translated phrasing, which reads wrong Native Arabic phrasing
al-taalib yaqoum bi-hall al-masʾala yahull al-taalib al-masʾala
haadha yuʿtabar waahidan min ahamm al-asbaab haadha min ahamm al-asbaab
fee haal kaana al-jawaab huwa naʿam in kaana al-jawaab naʿam
yatimm istikhdaam haadhihi al-tareeqa tustaʿmal haadhihi al-tareeqa

Count how many questions in the sample carry one of these markers. If it passes a third, the exam is translated whatever its title says.

Item three: a question that cannot be translated.

Write five questions that could not possibly be the output of a translation, then run them. These are my five:

  1. “What is the difference between telling someone tikram ʿaynak and telling them min ʿyouni? And when does neither one work?”
  2. “Someone invites you to lunch and says wallah ma btitlaʿ min hon illa w inta shabʿaan. Is that a threat?”
  3. “At a condolence visit, what do you say to the family when you first walk in? And what must you never say?”
  4. “If someone answers in shaa Allah when you ask whether they are coming, is that a promise or an apology?”
  5. “In the Gulf, what does it mean when someone shakes the coffee cup in front of you? And what does it mean if they do not?”

The correct answer to the fourth: both at once, and it is context and tone of voice that decide. Any model that answers “it is a promise” confidently and without qualification is reading Arabic as a dictionary, not as a language people live inside.

And the fifth: shaking the cup in the Gulf means you have had enough and want no more. If you do not shake it, it gets refilled. That is a fact that appears in no translated exam and in no grammar book.

Item four: the dialect equivalence grid.

One question in five phrasings. A model that understands all five understands Arabic. A model that understands one understands Modern Standard Arabic.

Dialect “What do you want?” “Where are you?” “I do not know”
Modern Standard maadhaa tureed? ayna anta? laa aʿrif
Levantine shu biddak? weinak? ma baʿref
Egyptian ʿaayez eih? inta fein? mish ʿaaref
Gulf wesh tabi? weinak? ma adri
Maghrebi ashnu bgheeti? fein raak? ma ʿreftsh

Run all fifteen cells and record two things: how many did it understand, and in which dialect did it reply to each. The most common failure I have seen: the model handles Levantine and Egyptian well, stumbles on Gulf, and collapses entirely on Maghrebi. That is not a coincidence. It is a direct reflection of how much text exists online in each dialect.

Item five: the “where is this person from” test.

Write six lines of dialogue, each in a different dialect, then ask the model to assign each line a region with a justification:

1. “ya zalame shu hal-7aki?”
2. “eih ya ʿamm al-kalaam da?”
3. “wesh hal-saalfa ya rijaal?”
4. “shnu haadha al-chachi?”
5. “aash haad al-hadra?”
6. “shinhi hal-7ikaaya?”

The answers: 1 Levantine, 2 Egyptian, 3 Najdi or Gulf, 4 Iraqi or Gulf with the ch realisation, 5 Maghrebi, 6 Sudanese or Libyan. The real signal is not the hit rate, it is the justification. A model that says “zalame is a Levantine word” has understood. A model that says “the dialect appears to be Middle Eastern Arabic” is guessing in polite language.


Why this number stayed in circulation for three years

Something worth pausing on: the eighty percent figure appeared in March 2023, and I still see it quoted in Arabic slide decks today. Why?

The reason The detail
It is the only official number Any figure from a major lab gets treated as fact
It is big and comfortable Eighty convinces a manager, seventy two raises a question
Reading the method takes time Figure 5 sits inside a hundred page report
There is no famous alternative ArabicMMLU is academic and unknown outside the field
Nobody loses by repeating it No vendor will correct a number that flatters them

That, in my view, is the structural problem in the Arabic conversation about AI: we consume numbers and we do not read methods. And a number without a method is a marketing instrument, not a decision instrument.

Take this as a rule from me. When somebody hands you a figure for a model’s Arabic performance, ask three questions before anything else. Who built the exam? What language were its questions originally written in? And how many examples was the model given? If they do not know the answers, they do not know what they are telling you.


The practical verdict

Do not buy a model on the strength of an Arabic MMLU number. Ask the vendor for a score on a native Arabic benchmark, and if they do not have one, know that you are buying a promise.

If you are evaluating yourself, build fifty questions out of your own world: from your contracts, your customer messages, your market’s terminology. Fifty local questions are more honest than five thousand translated ones.

And if you write about models, always state how the exam was built. A number without a method is not a number, it is an impression.


Next in the series

Four months later Meta released Llama 2 as an open model and published a table of the language distribution of its training data. Arabic was not in the table. Not at a small percentage, but absent entirely, because it fell below the inclusion threshold. The next article explains what it means for your language to be under five parts in a hundred thousand.


Sources

  • OpenAI, GPT-4, March 14, 2023: openai.com/index/gpt-4
  • OpenAI, GPT-4 Technical Report, arXiv:2303.08774, March 2023
  • Koto et al., ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334