The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Total training tokens | 395 billion | The model’s overall scale | Jais paper, arXiv:2308.16149, August 2023 |
| Original Arabic | 55 billion | How much genuine Arabic went in | Paper, Table 3 |
| Arabic after adding translated text | 72 billion | So 17 billion of it is translated | Same table |
| Arabic after 1.6x upsampling | 116 billion | The number that always gets announced | Same table |
| Arabic share of the total | 29% | After upsampling, not before | Same table |
| English | 232 billion, 59% | The largest language inside an Arabic model | Same table |
| Programming code | 46 billion, 12% | The remainder | Same table |
| Jais 13B base, Arabic benchmark average | 46.5% | Zero-shot | Paper, Table 9 |
| Jais-chat 13B | 48.4% | 5.5 points above BLOOMz | Same table |
| Best prior competitor, BLOOMz 7.1B | 42.9% | The reference | Same table |
| Jais-chat 13B on ArabicMMLU, a native Arabic exam | 54.8% | Against 72.5% for GPT-4 | Koto et al., February 2024 |
| Jais-chat 30B on ArabicMMLU | 62.3% | The best open model of its time | Same source |
The second row and the fourth row are the story: 55 billion of original Arabic became 116 billion in the announcement. The road between them is worth explaining.
Anatomy of the number
Upsampling is not cheating. It is a known, legitimate training technique and the paper states it plainly. But 116 billion is not 116 billion of distinct Arabic text. It is 72 billion read roughly one and a half times. And the difference matters, because repeating text adds no new knowledge. It deepens what is already there.
Model card
| Item | Detail |
|---|---|
| Model | Jais, at 13 and 30 billion parameters |
| Institutions | Inception, a G42 company; Mohamed bin Zayed University of Artificial Intelligence; Cerebras |
| Release date | August 30, 2023 |
| Paper | arXiv:2308.16149 |
| Compute | Condor Galaxy 1 |
| Licence | Open |
| First of its kind | The first large model deliberately built for Arabic |
| Evaluation benchmarks | Translated or adapted from English |
The ARAB-LENS reading: a real achievement and an incomplete story
Let us be clear: Jais was the most important event in digital Arabic in ten years. Before it there was no option. After it there was one. I will not undercut that by a word.
But three things in the numbers deserve to be said out loud.
First: half the model is English.
Read the table again. 232 billion English tokens against 116 billion Arabic ones. The “Arabic” model contains roughly twice as much English as Arabic. And that is a correct engineering decision, not a mistake. Reasoning, general knowledge and mathematics come from English, and you cannot build a capable model on fifty five billion Arabic tokens alone.
But when you read a headline that says “an Arabic model”, you picture an Arabic model. The reality is a bilingual model, mostly English, optimised for Arabic. And that is not the story that reached the public.
Second: seventeen billion of it is translated.
From 55 to 72 billion, the difference is Arabic machine-translated out of English. Which means roughly a quarter of the model’s original Arabic is not Arabic that an Arab wrote. I know why they did it: high-quality Arabic on the internet is scarce, and translation fills the gap.
The cost? The model learns patterns nobody says. It learns yatimm istikhdaam instead of tustaʿmal, and fee haal kaana instead of in kaana. Then it hands us Arabic that is grammatically sound and rhythmically foreign. This is what I call translation Arabic, and you will recognise it the moment you read it.
Third, and most serious: the exam is translated too.
Look at the evaluation suite: MMLU, HellaSwag, PIQA, BoolQ, ARC, WinoGrande, TruthfulQA. Every one is an English exam translated or adapted. Which means the loop closed: trained partly on translated Arabic, evaluated on translated Arabic, and the result announced as an Arabic achievement.
The proof that this matters arrived six months later. When Jais was tested on ArabicMMLU, an exam built from real Arab school questions, Jais-chat 13B scored 54.8%. That is a very good number for an open model at that time, and Jais-chat 30B scored 62.3% and was the best open model of any kind on that exam. So the model held up against a real Arabic test, and it deserves the credit.
My conclusion: Jais succeeded in spite of its measurement methodology, not because of it. If the evaluation had been natively Arabic from the start, the team would have seen its weak points earlier and fixed them.
From the test notebook: spotting “translation Arabic” in a model’s output
These are the items I run on any Arabic model to find out where it learned.
Item one: the fingerprint list.
Ask the model for a three hundred word article on any subject, then hunt for these:
| The fingerprint | The natural Arabic alternative | Why it appears |
|---|---|---|
| yatimm + verbal noun | The direct passive | A translation of “is being” |
| qaama bi + verbal noun | The plain verb | A translation of “did” |
| waahid min ahamm | min ahamm | A translation of “one of the most” |
| fee haal kaana | in kaana, idhaa | A translation of “in case” |
| bi-shakl + adjective | An adverb or accusative | A translation of the “-ly” ending |
| ʿalaa al-raghm min haqeeqat anna | maʿa anna, raghma anna | “Despite the fact that” |
| haadhaa yuʿtabar | haadhaa | “Is considered” |
Count the fingerprints per hundred words. In my experience with models built on translated data the rate passes six per hundred. In natively written Arabic it falls below two.
Item two: the proverb test.
Ask the model to complete five proverbs. Translation fails here, because proverbs do not translate:
| The opening | The correct completion |
|---|---|
| “illi eedo bil-mayy” | “mish mitl illi eedo bil-naar” |
| “al-ʾird bi-ʿein immo” | “ghazaal” |
| “tajree al-riyaah” | “bimaa laa tashtahee al-sufun” |
| “ya daakhil bein al-basale w ʾishritha” | “ma binoubak illa reehitha” |
| “illi ma byaʿref al-saʾr” | “byishweeh” |
A model that invents endings which are logical but do not exist has learned the structure of Arabic without having lived in it.
Item three: the greeting and sign-off test.
A simple and very sharp item. Ask: “Write me a short message to a work colleague apologising that I will not be able to attend tomorrow’s meeting.”
Watch four things:
| Element | Good sign | Bad sign |
|---|---|---|
| Opening | “marhaba” or “sabaah al-khair” | “ʿazeezee al-zameel al-muhtaram” |
| Apology | “baʿtizer minnak” | “awadd an uʿrib ʿan iʿtidhaaree” |
| Reason | Direct and short | A full paragraph of justification |
| Closing | “baʿtemed ʿaleik” or “minshoufak” | “wa tafaddaloo bi-qubool faaʾiq al-ihtiraam” |
The right-hand column is not a language error. It is a register error. A WhatsApp message to a colleague is not an official letter to a ministry. A model that cannot tell them apart learned Arabic from official and translated documents, which is the most abundant kind online.
Item four: measuring MSA leakage across six turns.
Start a conversation in Levantine and keep it going for six turns, recording the turn at which the model flips to Modern Standard Arabic:
Turn 1: “marhaba, keefak?”
Turn 2: “biddi asʾalak ʿan shaghle bil-shughl.”
Turn 3: “mudeeri talab minni taqreer w ana ma fhimt shu biddo bil-zabt.”
Turn 4: “laʾ, huwwe ma haka li. bas ʾaal jahhezli shi murattab.”
Turn 5: “tayyeb shu raʾyak aʿmal?”
Turn 6: “tamaam, bas khalleeha ʾaseere laʾenno ma ʿindi waʾt.”
The pattern I have watched repeat: turns one and two in Levantine, the third carrying formal words, the fourth fully formal. Record the turn number. It is a very useful and very fast metric, and I use it across this entire series.
Item five: the internal bilingualism test.
A bilingual model may think in one language and answer in another. Ask for a problem requiring steps, and ask for the steps in Arabic:
“I have a shop. I sold 17 items at 23 dinars each and paid 40 dinars for delivery. What is my net income? Write the steps in Arabic.”
Watch: are the steps genuinely Arabic, or numbers and English words? Did it keep “dinar” or convert it? Did it order the digits correctly? A model that gives you the right answer with half-English steps is thinking in English and translating.
Item six: the proper noun test.
Ask for a paragraph about an Arab figure who is not globally famous, such as the poet al-Mutanabbi, or Ibn al-Nafis, or Ghassan Kanafani. Then ask for the same paragraph in English. Compare the number of correct facts in each. If the English version is richer, the model stores its knowledge of your culture in somebody else’s language.
What Jais teaches us against what came before
| Criterion | Before Jais | After it |
|---|---|---|
| Best open Arabic model | BLOOMz at 42.9% | Jais-chat at 48.4% |
| An open Arabic option at serious scale | No | Yes |
| Transparency about Arabic training data | None | A detailed table |
| Buildability for regional teams | Difficult | Direct |
| Evaluation on native Arabic benchmarks | No | Arrived later with ArabicMMLU |
The third row is what I value most. Table 3 of the Jais paper breaks it down: original, translated, upsampled. That is transparency I have not found in most of what I read afterwards, including work two years newer.
And I will close with an observation I think matters for the region: transparency about weakness is more useful than an announcement about strength. A team that writes “a quarter of our Arabic is translated” gives the researchers who follow them a foundation to build on. A team that publishes one shiny number gives them a press headline.
The practical verdict
If you are choosing between a dedicated Arabic model and a global one, the rule I have arrived at: the dedicated Arabic model is better at register and cultural context, and the large global model is better at knowledge and reasoning. Choose by your task, not by the name.
If you are building Arabic training data, do not fill the gap with machine translation. Five billion tokens of native Arabic are worth more than twenty billion half of which is translated. And that is not only my opinion. The AraT5 paper proved it with numbers in 2022.
If you are reading an announcement about an Arabic model, ask for three numbers: how much original Arabic, how much translated, and how many times it was repeated. The announced figure alone tells you nothing.
Next in the series
One week after Jais, the UAE also announced Falcon 180B, the largest open model in the world at that point, trained on three and a half trillion tokens. The next article goes looking for Arabic in the model card. And does not find it.
Sources
- Sengupta et al., Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models, arXiv:2308.16149, August 30, 2023
- Cerebras, G42 and MBZUAI press release, August 2023: cerebras.ai
- Koto et al., ArabicMMLU, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334
- Huang et al., AceGPT, NAACL 2024: aclanthology.org/2024.naacl-long.450