Half of Its Arabic Is Machine-Translated: Reading ALLaM and What “National Model” Means

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Paper released: July 22, 2024 | Published: August 12, 2024

The numbers first

Metric Result What it measures Source
Total Arabic curated 540 billion tokens The size of the Arabic corpus ALLaM paper, arXiv:2407.15390
Of which natural Arabic 270 billion Half the total Same paper, quoted verbatim
Of which machine-translated Arabic 270 billion The other half Same paper
English available for pretraining 4 trillion tokens The foundation Same paper
The training recipe 4 trillion English, then 1.2 trillion mixed Arabic and English 5.2 trillion in total Model card and paper
Starting point Llama 2 weights The base used Same paper
ALLaM-Instruct 7B on Arabic MMLU 66.9 Against 53.98 for Llama 3-Instruct 8B Paper, Table 4
On ACVA 80.33 Against 75.21 for Llama 3-Instruct 8B Same table
On araSwag 49.28 Against 33.99 for Llama 3-Instruct 8B Same table
On ETEC 52.7 Against 44.32 for Llama 3-Instruct 8B Same table
On araMath 36.4 Against 37.9 for AceGPT-Chat 7B, the only one it loses Same table
Model card, average across 12 benchmarks ALLaM-7B-Instruct 64.42 A later and different evaluation Card humain-ai/ALLaM-7B-Instruct-preview
On the same table, Qwen2.5-72B-Instruct 76.91 A general Chinese model, 12.5 points higher Same card
On the same table, Qwen2.5-7B-Instruct 60.55 The comparable size, four points lower Same card

Two numbers deserve a pause: half the Arabic corpus is translated, and a general Chinese model outscores the Saudi model on Arabic, in figures published on the Saudi model’s own card.


Anatomy of the Arabic corpus

540 billion Arabic tokens, as the paper describes them
Natural Arabic
270 billion
Translated Arabic
270 billion

Exactly half and half. This is the first time I have seen a team state that proportion with this clarity, and I respect the transparency even when I do not like the number.


Model card

Item Detail
Model ALLaM, at 7, 13 and 70 billion parameters
Institution The National Center for AI, Saudi Data and AI Authority
Paper date July 22, 2024
7B weights available February 13, 2025
Authority announcement March 6, 2025
Base Llama 2 weights, with continued pretraining
Arabic corpus 540 billion tokens, half translated
Benchmarks araSwag, ACVA, Arabic MMLU, Exams, ETEC, araTruthfulQA, araMath

The ARAB-LENS reading: three things, two of them deserving credit

First, and the most positive: the numbers are very good.

Look at Table 4 carefully. ALLaM-Instruct at seven billion parameters scored 66.9 on Arabic MMLU, against 53.98 for Llama 3-Instruct at eight billion. Thirteen points of difference, in favour of the smaller model. On araSwag: 49.28 against 33.99, fifteen points. On ETEC: 52.7 against 44.32.

That is a real achievement and it should be said plainly. An Arabic model at seven billion parameters beats Meta’s newer model of comparable size, in Arabic, by margins that are large rather than marginal. And it proves the hypothesis that keeps recurring in this series: specialisation works.

Second: transparency about translation.

The paper states in so many words that 270 of 540 billion Arabic tokens are machine-translated. Nobody had to infer it. The team wrote it.

I value that a great deal, because in every article of this series I had to dig for this information. Here it was written down.

But appreciation does not preclude criticism: half the corpus is translated. And we saw in the AraT5 article, with numbers from ACL 2022, that less Arabic carefully chosen beats more Arabic of lower quality. And we saw in the AlGhafa article that human reviewers accepted only 58% of machine-translated questions.

So the legitimate question is: what if the model had been trained on the 270 billion natural tokens alone, with the extra effort spent on cleaning rather than translating? I do not know the answer, and nobody does, because the experiment was not run. But the question deserves to be put to every Arab team building today.

Third, and this must be read coolly: a general Chinese model outscores it.

On the model’s Hugging Face card there is an evaluation table across twelve benchmarks. ALLaM-7B-Instruct averages 64.42. Qwen2.5-72B-Instruct averages 76.91. Twelve and a half points in favour of a general purpose Chinese model.

Before this is misread, three necessary observations:

The observation The detail
The sizes differ 72 billion against 7. That comparison is unfair by nature
At comparable size Qwen2.5-7B scored 60.55, which is below ALLaM by four points
The table is SDAIA’s own publication The authority published a number that does not flatter it, and that counts in its favour

So the correct reading: ALLaM wins in its weight class and loses to something ten times its size. That is not a defeat. That is physics.

But the strategic lesson stands, and it is where I want to end: when a large general model reaches the market, it overtakes the small specialised model even inside its speciality. And the region cannot compete on scale today. So where does it compete?

My answer after twenty articles: in what scale cannot buy. In dialects, in local data, in benchmarks, in cultural context, in specialised applications. Those things do not arrive with more processing units. They arrive with slow fieldwork.


From the test notebook: the final test in this set

I will leave here the shortest and harshest test I have, the one I run first on any new Arabic model. Ten items, each a single line, scored zero to two.

# The item The correct answer in brief
1 In “3am baktob la3ammi”, what is each “3am” doing? The first is a progressive marker, the second is the word for uncle
2 “al-kutub jadeeda” or “al-kutub judud”? jadeeda, because books are non-human
3 “thalaath sayyaaraat” or “thalaathat sayyaaraat”? thalaath, because sayyaara is feminine so the number goes masculine
4 “kitaab al-taalib” or “al-kitaab al-taalib”? kitaab al-taalib, the first noun of a construct takes no article
5 Someone said “in shaa Allah” when asked if they are coming. Promise or apology? Both, and context decides
6 Your manager said “khalleena nshoof” to a raise request. What does it mean? Usually a deferred and polite refusal
7 Your friend’s uncle died and you never met him. Do you attend the condolence? Yes, the duty is towards your friend
8 What is the difference between ʿAmmaan and ʿUmaan? Gemination and vowel, and two different countries
9 “gaal”, “ʾaal”, “kaal”, “qaal”: what do they share? Four realizations of the letter qaf
10 In the Gulf, shaking the coffee cup means what? I have had enough, do not pour more

How to read the score:

Total out of 20 The reading
0 to 6 It writes Arabic, it does not know it
7 to 12 It knows formal Arabic, it does not know people
13 to 16 Good, and usable in a product with review
17 to 20 Rare, and I have not yet seen anything reach 20

Items 1 to 4 are grammatical, 5 to 7 pragmatic, 8 to 10 dialectal and cultural. And the model that passes the first four and falls on the remaining six is the most common kind, and it is exactly the model that looks excellent in a demo and lets you down with your first real customer.


The three Arab paths compared

This set has seen three major Arab projects. I place them side by side:

Item Jais, 2023 AceGPT, 2023 ALLaM, 2024
Institution UAE Academic, Saudi and Chinese Saudi Arabia
Base From scratch On top of Llama 2 On top of Llama 2
Original Arabic 55 billion tokens Not precisely declared 270 billion tokens
Translated Arabic 17 billion Not declared 270 billion
Transparency High High High
Strength The first serious Arabic model The localisation methodology The best numbers in its class

Note the row before strength: transparency is high in all three. That is the best habit in the region and I hope it lasts. Arab projects publish far more detail about their data than the global companies do.

I believe the reason is that these are academic or national projects at heart rather than purely commercial products. And that is a situation worth protecting, because the moment these models become purely commercial products, the published detail will shrink.

The last item in the test notebook: the vendor question.

When any vendor offers you an Arabic model, ask these five and write down the answers:

The question Why
How many original Arabic tokens in training? Separates the serious from the marketed
How many of those are machine-translated? Exposes corpus quality
Which native Arabic benchmark did you measure on? Not a translated one
What is its dialect performance, broken down? Most vendors do not know
Who measured, you or an independent party? Determines the weight of the number

Anyone who answers all five with numbers deserves the rest of the conversation. Anyone who answers three, proceed with caution. Anyone who answers none is not selling you a model. They are selling you a slide deck.


The practical verdict

The specialised model wins in its weight class, so choose it if your budget is limited. ALLaM at seven billion beats Qwen2.5 at seven billion and Llama 3 at eight, in Arabic. That is clear and published.

And if your budget allows a large general model, it will usually win. That is also clear and published, in the same document.

And the decision is not only technical. Hosting inside the region, compliance with local data law, and support in your language are legitimate reasons to choose a regional model even if it loses points on a benchmark.


Closing this set

This set of twenty articles began in May 2022, when Arabic had models at hundreds of millions of parameters, native benchmarks, and a methodology that separated formal Arabic from dialect. It ends in August 2024, when Arabic has models at tens of billions, better benchmarks, and more published numbers.

The thread running through all twenty is one: announcement is always faster than measurement, Arabic measurement arrives from outside and arrives late, and whoever does not measure for themselves is buying promises.

And the good news is that this is changing. A leaderboard that corrects itself in public, a native Arabic benchmark of fourteen thousand questions, and Saudi, Emirati and Egyptian teams publishing numbers. Slow is not the same as stalled. It is a beginning.


Sources

  • Bari et al., ALLaM: Large Language Models for Arabic and English, arXiv:2407.15390, July 22, 2024
  • Model card ALLaM-7B-Instruct-preview on Hugging Face
  • Saudi Data and AI Authority announcement, Saudi Press Agency, March 6, 2025
  • Nagoudi, Elmadany and Abdul-Mageed, AraT5, ACL 2022, for the curation against scale comparison
  • Almazrouei et al., AlGhafa, ArabicNLP 2023, for the machine translation acceptance rate