The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Total Arabic curated | 540 billion tokens | The size of the Arabic corpus | ALLaM paper, arXiv:2407.15390 |
| Of which natural Arabic | 270 billion | Half the total | Same paper, quoted verbatim |
| Of which machine-translated Arabic | 270 billion | The other half | Same paper |
| English available for pretraining | 4 trillion tokens | The foundation | Same paper |
| The training recipe | 4 trillion English, then 1.2 trillion mixed Arabic and English | 5.2 trillion in total | Model card and paper |
| Starting point | Llama 2 weights | The base used | Same paper |
| ALLaM-Instruct 7B on Arabic MMLU | 66.9 | Against 53.98 for Llama 3-Instruct 8B | Paper, Table 4 |
| On ACVA | 80.33 | Against 75.21 for Llama 3-Instruct 8B | Same table |
| On araSwag | 49.28 | Against 33.99 for Llama 3-Instruct 8B | Same table |
| On ETEC | 52.7 | Against 44.32 for Llama 3-Instruct 8B | Same table |
| On araMath | 36.4 | Against 37.9 for AceGPT-Chat 7B, the only one it loses | Same table |
| Model card, average across 12 benchmarks | ALLaM-7B-Instruct 64.42 | A later and different evaluation | Card humain-ai/ALLaM-7B-Instruct-preview |
| On the same table, Qwen2.5-72B-Instruct | 76.91 | A general Chinese model, 12.5 points higher | Same card |
| On the same table, Qwen2.5-7B-Instruct | 60.55 | The comparable size, four points lower | Same card |
Two numbers deserve a pause: half the Arabic corpus is translated, and a general Chinese model outscores the Saudi model on Arabic, in figures published on the Saudi model’s own card.
Anatomy of the Arabic corpus
Exactly half and half. This is the first time I have seen a team state that proportion with this clarity, and I respect the transparency even when I do not like the number.
Model card
| Item | Detail |
|---|---|
| Model | ALLaM, at 7, 13 and 70 billion parameters |
| Institution | The National Center for AI, Saudi Data and AI Authority |
| Paper date | July 22, 2024 |
| 7B weights available | February 13, 2025 |
| Authority announcement | March 6, 2025 |
| Base | Llama 2 weights, with continued pretraining |
| Arabic corpus | 540 billion tokens, half translated |
| Benchmarks | araSwag, ACVA, Arabic MMLU, Exams, ETEC, araTruthfulQA, araMath |
The ARAB-LENS reading: three things, two of them deserving credit
First, and the most positive: the numbers are very good.
Look at Table 4 carefully. ALLaM-Instruct at seven billion parameters scored 66.9 on Arabic MMLU, against 53.98 for Llama 3-Instruct at eight billion. Thirteen points of difference, in favour of the smaller model. On araSwag: 49.28 against 33.99, fifteen points. On ETEC: 52.7 against 44.32.
That is a real achievement and it should be said plainly. An Arabic model at seven billion parameters beats Meta’s newer model of comparable size, in Arabic, by margins that are large rather than marginal. And it proves the hypothesis that keeps recurring in this series: specialisation works.
Second: transparency about translation.
The paper states in so many words that 270 of 540 billion Arabic tokens are machine-translated. Nobody had to infer it. The team wrote it.
I value that a great deal, because in every article of this series I had to dig for this information. Here it was written down.
But appreciation does not preclude criticism: half the corpus is translated. And we saw in the AraT5 article, with numbers from ACL 2022, that less Arabic carefully chosen beats more Arabic of lower quality. And we saw in the AlGhafa article that human reviewers accepted only 58% of machine-translated questions.
So the legitimate question is: what if the model had been trained on the 270 billion natural tokens alone, with the extra effort spent on cleaning rather than translating? I do not know the answer, and nobody does, because the experiment was not run. But the question deserves to be put to every Arab team building today.
Third, and this must be read coolly: a general Chinese model outscores it.
On the model’s Hugging Face card there is an evaluation table across twelve benchmarks. ALLaM-7B-Instruct averages 64.42. Qwen2.5-72B-Instruct averages 76.91. Twelve and a half points in favour of a general purpose Chinese model.
Before this is misread, three necessary observations:
| The observation | The detail |
|---|---|
| The sizes differ | 72 billion against 7. That comparison is unfair by nature |
| At comparable size | Qwen2.5-7B scored 60.55, which is below ALLaM by four points |
| The table is SDAIA’s own publication | The authority published a number that does not flatter it, and that counts in its favour |
So the correct reading: ALLaM wins in its weight class and loses to something ten times its size. That is not a defeat. That is physics.
But the strategic lesson stands, and it is where I want to end: when a large general model reaches the market, it overtakes the small specialised model even inside its speciality. And the region cannot compete on scale today. So where does it compete?
My answer after twenty articles: in what scale cannot buy. In dialects, in local data, in benchmarks, in cultural context, in specialised applications. Those things do not arrive with more processing units. They arrive with slow fieldwork.
From the test notebook: the final test in this set
I will leave here the shortest and harshest test I have, the one I run first on any new Arabic model. Ten items, each a single line, scored zero to two.
| # | The item | The correct answer in brief |
|---|---|---|
| 1 | In “3am baktob la3ammi”, what is each “3am” doing? | The first is a progressive marker, the second is the word for uncle |
| 2 | “al-kutub jadeeda” or “al-kutub judud”? | jadeeda, because books are non-human |
| 3 | “thalaath sayyaaraat” or “thalaathat sayyaaraat”? | thalaath, because sayyaara is feminine so the number goes masculine |
| 4 | “kitaab al-taalib” or “al-kitaab al-taalib”? | kitaab al-taalib, the first noun of a construct takes no article |
| 5 | Someone said “in shaa Allah” when asked if they are coming. Promise or apology? | Both, and context decides |
| 6 | Your manager said “khalleena nshoof” to a raise request. What does it mean? | Usually a deferred and polite refusal |
| 7 | Your friend’s uncle died and you never met him. Do you attend the condolence? | Yes, the duty is towards your friend |
| 8 | What is the difference between ʿAmmaan and ʿUmaan? | Gemination and vowel, and two different countries |
| 9 | “gaal”, “ʾaal”, “kaal”, “qaal”: what do they share? | Four realizations of the letter qaf |
| 10 | In the Gulf, shaking the coffee cup means what? | I have had enough, do not pour more |
How to read the score:
| Total out of 20 | The reading |
|---|---|
| 0 to 6 | It writes Arabic, it does not know it |
| 7 to 12 | It knows formal Arabic, it does not know people |
| 13 to 16 | Good, and usable in a product with review |
| 17 to 20 | Rare, and I have not yet seen anything reach 20 |
Items 1 to 4 are grammatical, 5 to 7 pragmatic, 8 to 10 dialectal and cultural. And the model that passes the first four and falls on the remaining six is the most common kind, and it is exactly the model that looks excellent in a demo and lets you down with your first real customer.
The three Arab paths compared
This set has seen three major Arab projects. I place them side by side:
| Item | Jais, 2023 | AceGPT, 2023 | ALLaM, 2024 |
|---|---|---|---|
| Institution | UAE | Academic, Saudi and Chinese | Saudi Arabia |
| Base | From scratch | On top of Llama 2 | On top of Llama 2 |
| Original Arabic | 55 billion tokens | Not precisely declared | 270 billion tokens |
| Translated Arabic | 17 billion | Not declared | 270 billion |
| Transparency | High | High | High |
| Strength | The first serious Arabic model | The localisation methodology | The best numbers in its class |
Note the row before strength: transparency is high in all three. That is the best habit in the region and I hope it lasts. Arab projects publish far more detail about their data than the global companies do.
I believe the reason is that these are academic or national projects at heart rather than purely commercial products. And that is a situation worth protecting, because the moment these models become purely commercial products, the published detail will shrink.
The last item in the test notebook: the vendor question.
When any vendor offers you an Arabic model, ask these five and write down the answers:
| The question | Why |
|---|---|
| How many original Arabic tokens in training? | Separates the serious from the marketed |
| How many of those are machine-translated? | Exposes corpus quality |
| Which native Arabic benchmark did you measure on? | Not a translated one |
| What is its dialect performance, broken down? | Most vendors do not know |
| Who measured, you or an independent party? | Determines the weight of the number |
Anyone who answers all five with numbers deserves the rest of the conversation. Anyone who answers three, proceed with caution. Anyone who answers none is not selling you a model. They are selling you a slide deck.
The practical verdict
The specialised model wins in its weight class, so choose it if your budget is limited. ALLaM at seven billion beats Qwen2.5 at seven billion and Llama 3 at eight, in Arabic. That is clear and published.
And if your budget allows a large general model, it will usually win. That is also clear and published, in the same document.
And the decision is not only technical. Hosting inside the region, compliance with local data law, and support in your language are legitimate reasons to choose a regional model even if it loses points on a benchmark.
Closing this set
This set of twenty articles began in May 2022, when Arabic had models at hundreds of millions of parameters, native benchmarks, and a methodology that separated formal Arabic from dialect. It ends in August 2024, when Arabic has models at tens of billions, better benchmarks, and more published numbers.
The thread running through all twenty is one: announcement is always faster than measurement, Arabic measurement arrives from outside and arrives late, and whoever does not measure for themselves is buying promises.
And the good news is that this is changing. A leaderboard that corrects itself in public, a native Arabic benchmark of fourteen thousand questions, and Saudi, Emirati and Egyptian teams publishing numbers. Slow is not the same as stalled. It is a beginning.
Sources
- Bari et al., ALLaM: Large Language Models for Arabic and English, arXiv:2407.15390, July 22, 2024
- Model card ALLaM-7B-Instruct-preview on Hugging Face
- Saudi Data and AI Authority announcement, Saudi Press Agency, March 6, 2025
- Nagoudi, Elmadany and Abdul-Mageed, AraT5, ACL 2022, for the curation against scale comparison
- Almazrouei et al., AlGhafa, ArabicNLP 2023, for the machine translation acceptance rate