Less Than Five Parts in a Hundred Thousand: Arabic in Llama 2’s Training Data

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: July 18, 2023 | Published: August 1, 2023

The numbers first

Metric Result What it measures Source
Arabic share of Llama 2 pretraining data Not listed, meaning below 0.005% How much Arabic entered the model at all Llama 2 paper, Table 10, July 2023
English share 89.70% The scale of the dominance Same table
“Unknown” language share 8.38% The undocumented block Same table
Highest non-English language, German 0.17% The real ceiling for any other language Same table
Inclusion threshold for the table 0.005% Why Arabic vanished Same table
LLaMA2-7B on Arabic MMLU 29.47% Against 29.0% for random guessing AceGPT paper, NAACL 2024
LLaMA2-7B on Arabic EXAMs 23.48% Below random Same paper
LLaMA2-13B on ArabicMMLU 36.1% Against 72.5% for GPT-4 Koto et al., ArabicMMLU
Llama 2-Chat 7B on araSwag 24.44 The weakest figure in the entire comparison table ALLaM paper, July 2024

The sixth row is the one that should stop you: 29.47 against 29.0 for random guessing. Half a point above chance. That is not weak performance. That is the absence of performance.


The language table, as Meta published it

Language distribution of Llama 2 pretraining data
English
89.70%
Unknown
8.38%
German
0.17%
French
0.16%
Chinese
0.13%
Arabic (no row exists) below 0.005%

Note the last line. There is no bar because there is no number. Arabic, the language of roughly half a billion people, did not reach a threshold of five parts in a hundred thousand.


Model card

Item Detail
Model Llama 2, at 7, 13 and 70 billion parameters
Lab Meta
Release date July 18, 2023
Paper arXiv:2307.09288
Licence Open for commercial use with conditions
Arabic among the stated languages No mention
Meta’s own warning “A training corpus with a majority in English means that the model may not be suitable for use in other languages”

That last line matters and I will come back to it. Meta warned us. The paper says it in those words. And half the Arab world built its products on this model anyway, because it was free.


The ARAB-LENS reading: the zero nobody apologised for

I will say something that may sound harsh: Llama 2 did not fail at Arabic. Llama 2 never attempted Arabic. The difference between those two is the difference between a student who failed the exam and a student who never entered the hall.

Look at the table again. Arabic was not the only language to vanish. Hindi vanished. Bengali, Turkish, Persian, Hebrew, Indonesian, Swahili. The table holds roughly twenty six languages, all European or East Asian, and everything else is zero. This is not a decision against Arabic. It is a decision in favour of English alone, and Arabic is among the casualties.

Now the thing I want you to understand properly, because it explains a decade of Arab frustration with these models.

The 8.38% “unknown” block is the only place Arabic could have been hiding. Meta does not break it down. I cannot claim Arabic is inside it and I cannot claim it is not. What I can say is that even if every Arabic token in the model sits inside that block, it entered unclassified, uncleaned and unweighted. Random Arabic text that fell into the dough.

And that explains a behaviour I have watched repeat with every model built on Llama 2. The model knows the shape of the Arabic letter, knows a few hundred common words, and can assemble a sentence that looks Arabic. Then you ask it something that requires knowledge and it drops to guessing level. The form is there, the content is not. Which is exactly what 29.47 against 29.0 is telling you.

The final observation, and the one that matters most for the region: dozens of Arab companies and universities took Llama 2 and built on it, because the licence was open and the cost was zero. That was not a mistake. It was the only realistic option available. But the consequence is that an entire layer of Arabic products was built on a foundation that is less than 0.005% Arabic. And then we wonder why our products do not understand our dialects.

My honest view: the problem is not that Meta did not train on Arabic. The problem is that we built as though it had.


From the test notebook: how to expose a model whose Arabic is skin deep

Four items that separate “writes Arabic” from “knows Arabic”. Run them in order.

Item one: non-human plural agreement.

The Arabic rule: a plural of non-humans takes feminine singular agreement. A shallow model gets it wrong every time, because it learned agreement from English.

Correct Common error Why
al-kutub jadeeda al-kutub judud Books are non-human
al-sayyaaraat waaqifa al-sayyaaraat waaqifeen Cars are non-human
al-mashaakil katheera al-mashaakil kithaar Problems are non-human
al-muwazzafoon judud al-muwazzafoon jadeeda Employees are human, so here the reverse holds

Ask the model for a two hundred word paragraph containing at least five non-human plurals, then count the errors.

Item two: inverted number polarity.

In Arabic, numbers from three to ten take the opposite gender to the thing counted. No European language has this rule, so a shallow model simply does not have it.

Correct Error The rule
thalaathat kutub thalaath kutub kitaab is masculine, so the number goes feminine
thalaath sayyaaraat thalaathat sayyaaraat sayyaara is feminine, so the number goes masculine
khamsat rijaal khams rijaal rajul is masculine
khams nisaaʾ khamsat nisaaʾ imraʾa is feminine

Item three: the construct state.

Ask for a sentence containing a genitive construction, then inspect the definite article and the tanween:

Correct Error
kitaabu al-taalibi al-kitaabu al-taalibi
mudeeru al-sharika al-mudeeru al-sharika
baabu baytin qadeemin baabu al-baytin qadeemin

A model that puts the definite article on the first noun of a construct has not learned Arabic. It has learned to glue “al-” onto nouns.

Item four: the minimum dialect test.

Write in very simple Levantine and ask for a Levantine reply:

“Marhaba, ana biddi asʾal ʿan shaghle zgheere. Iza hada hajaz w baʿdein ma ija, bitrajjʿoulo al-masaari walla laʾ?”
(Hello, I want to ask about something small. If someone books and then does not show up, do you refund them or not?)

Record three things. Did it understand the question? Did it reply in Levantine or slide into Modern Standard Arabic? And did it keep biddi and masaari, or convert them to ureed and al-maal?

In models built on a weak foundation the reply is fully formal from the first sentence. That is not politeness. It is inability.

Item five: the hamza and its seat.

The rule for a medial hamza depends on the strength ranking of the vowels: kasra, then damma, then fatha, then sukun. A weak model writes the hamza at random because it memorised word shapes rather than learning a rule.

Correct Common error The stronger vowel
masʾool masʾuul written on the wrong seat damma
raʾees raees kasra
saʾala sears the wrong seat fatha
yaqraʾ yaqraa fatha then sukun
shayʾ shayi on a yaa seat sukun before the hamza
difʾ difʾi on a yaa seat sukun

Ask for a paragraph containing ten medial and final hamzas, then count. A model that misses more than three has not learned the rule. It memorised a list.

Item six: the local knowledge test.

These are questions that cannot be answered from translated English text:

The question What it measures
“What is the difference between Shobak and Karak?” Fine-grained Jordanian geography
“Why do people make that particular joke about the people of Homs?” Levantine folk humour
“What does it mean when someone tells you they are from the 48?” Palestinian context
“What is the difference between ʿaseeda and madeeda?” Gulf and Sudanese cooking
“Why is Friday the day off and not Sunday?” A cultural given

The last one is the trap. Any model trained on English text will say “the weekend” and attach it to Saturday and Sunday by default. A model that actually knows Arabic knows that the working week differs, and that it differs even between Arab countries.


Who built on this foundation, and what happened

This table is the most practically useful thing in the article, because it explains the map of Arabic models you will see in the articles that follow:

Model Built on The result
Jais Built from scratch, not on Llama Native Arabic, but weaker general knowledge
AceGPT Continued training on top of Llama 2 Large gain in dialogue, flat on knowledge
ALLaM Started from Llama 2 weights then continued A large jump, with half its Arabic data translated
Dozens of local models Fine-tuned on top of Llama 2 Surface improvement only

Notice the second and third rows. AceGPT and ALLaM both started from Llama 2. Two of the most important models the region has produced were built on a foundation that is under 0.005% Arabic. They improved it a great deal, and that is a real achievement. But the ceiling they hit is not the ceiling of their effort. It is the ceiling of the foundation.

In my view this explains a specific pattern: Arabic models built on Llama are good at style and bad at memory. They write beautiful Arabic and get a simple historical fact wrong. Because continued training teaches a model how to speak, and cannot plant in it what it never read in the first place.


The practical verdict

If you are choosing an open model for an Arabic product, ask one question before anything else: has the lab published the Arabic share of its training data? If the answer is no, assume it is near zero until proven otherwise.

If you are forced to build on a model with weak Arabic, continued pretraining on Arabic data helps, but do not expect a miracle. You are adding a layer on top of a foundation. You are not building a foundation.

And most important: do not measure your model by “does it write comprehensible Arabic”. That is the easiest test in the world and every model passes it. Measure it on plurals, numbers, construct state and dialect. That is where the difference shows.


Next in the series

One month after this launch the UAE announced Jais, the first large model deliberately built for Arabic, on a hundred and sixteen billion Arabic tokens. The next article explains that number, and also the detail it conceals: how 55 billion became 116 billion.


Sources

  • Meta, Meta and Microsoft Introduce the Next Generation of Llama, July 18, 2023: ai.meta.com/blog/llama-2
  • Touvron et al., Llama 2: Open Foundation and Fine-Tuned Chat Models, arXiv:2307.09288, Appendix A.4, Table 10
  • Huang et al., AceGPT: Localizing Large Language Models in Arabic, NAACL 2024: aclanthology.org/2024.naacl-long.450
  • Koto et al., ArabicMMLU, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334
  • Bari et al., ALLaM: Large Language Models for Arabic and English, arXiv:2407.15390, July 2024