A Hundred and Sixteen Billion Arabic Tokens: How Fifty Five Became a Hundred and Sixteen

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: August 30, 2023 | Published: September 13, 2023

The numbers first

Metric Result What it measures Source
Total training tokens 395 billion The model’s overall scale Jais paper, arXiv:2308.16149, August 2023
Original Arabic 55 billion How much genuine Arabic went in Paper, Table 3
Arabic after adding translated text 72 billion So 17 billion of it is translated Same table
Arabic after 1.6x upsampling 116 billion The number that always gets announced Same table
Arabic share of the total 29% After upsampling, not before Same table
English 232 billion, 59% The largest language inside an Arabic model Same table
Programming code 46 billion, 12% The remainder Same table
Jais 13B base, Arabic benchmark average 46.5% Zero-shot Paper, Table 9
Jais-chat 13B 48.4% 5.5 points above BLOOMz Same table
Best prior competitor, BLOOMz 7.1B 42.9% The reference Same table
Jais-chat 13B on ArabicMMLU, a native Arabic exam 54.8% Against 72.5% for GPT-4 Koto et al., February 2024
Jais-chat 30B on ArabicMMLU 62.3% The best open model of its time Same source

The second row and the fourth row are the story: 55 billion of original Arabic became 116 billion in the announcement. The road between them is worth explaining.


Anatomy of the number

Where the hundred and sixteen billion came from
Original Arabic collected
55 billion
+ machine-translated Arabic
17 billion
= the real total
72 billion
x upsampled 1.6 times
116 billion announced

Upsampling is not cheating. It is a known, legitimate training technique and the paper states it plainly. But 116 billion is not 116 billion of distinct Arabic text. It is 72 billion read roughly one and a half times. And the difference matters, because repeating text adds no new knowledge. It deepens what is already there.


Model card

Item Detail
Model Jais, at 13 and 30 billion parameters
Institutions Inception, a G42 company; Mohamed bin Zayed University of Artificial Intelligence; Cerebras
Release date August 30, 2023
Paper arXiv:2308.16149
Compute Condor Galaxy 1
Licence Open
First of its kind The first large model deliberately built for Arabic
Evaluation benchmarks Translated or adapted from English

The ARAB-LENS reading: a real achievement and an incomplete story

Let us be clear: Jais was the most important event in digital Arabic in ten years. Before it there was no option. After it there was one. I will not undercut that by a word.

But three things in the numbers deserve to be said out loud.

First: half the model is English.

Read the table again. 232 billion English tokens against 116 billion Arabic ones. The “Arabic” model contains roughly twice as much English as Arabic. And that is a correct engineering decision, not a mistake. Reasoning, general knowledge and mathematics come from English, and you cannot build a capable model on fifty five billion Arabic tokens alone.

But when you read a headline that says “an Arabic model”, you picture an Arabic model. The reality is a bilingual model, mostly English, optimised for Arabic. And that is not the story that reached the public.

Second: seventeen billion of it is translated.

From 55 to 72 billion, the difference is Arabic machine-translated out of English. Which means roughly a quarter of the model’s original Arabic is not Arabic that an Arab wrote. I know why they did it: high-quality Arabic on the internet is scarce, and translation fills the gap.

The cost? The model learns patterns nobody says. It learns yatimm istikhdaam instead of tustaʿmal, and fee haal kaana instead of in kaana. Then it hands us Arabic that is grammatically sound and rhythmically foreign. This is what I call translation Arabic, and you will recognise it the moment you read it.

Third, and most serious: the exam is translated too.

Look at the evaluation suite: MMLU, HellaSwag, PIQA, BoolQ, ARC, WinoGrande, TruthfulQA. Every one is an English exam translated or adapted. Which means the loop closed: trained partly on translated Arabic, evaluated on translated Arabic, and the result announced as an Arabic achievement.

The proof that this matters arrived six months later. When Jais was tested on ArabicMMLU, an exam built from real Arab school questions, Jais-chat 13B scored 54.8%. That is a very good number for an open model at that time, and Jais-chat 30B scored 62.3% and was the best open model of any kind on that exam. So the model held up against a real Arabic test, and it deserves the credit.

My conclusion: Jais succeeded in spite of its measurement methodology, not because of it. If the evaluation had been natively Arabic from the start, the team would have seen its weak points earlier and fixed them.


From the test notebook: spotting “translation Arabic” in a model’s output

These are the items I run on any Arabic model to find out where it learned.

Item one: the fingerprint list.

Ask the model for a three hundred word article on any subject, then hunt for these:

The fingerprint The natural Arabic alternative Why it appears
yatimm + verbal noun The direct passive A translation of “is being”
qaama bi + verbal noun The plain verb A translation of “did”
waahid min ahamm min ahamm A translation of “one of the most”
fee haal kaana in kaana, idhaa A translation of “in case”
bi-shakl + adjective An adverb or accusative A translation of the “-ly” ending
ʿalaa al-raghm min haqeeqat anna maʿa anna, raghma anna “Despite the fact that”
haadhaa yuʿtabar haadhaa “Is considered”

Count the fingerprints per hundred words. In my experience with models built on translated data the rate passes six per hundred. In natively written Arabic it falls below two.

Item two: the proverb test.

Ask the model to complete five proverbs. Translation fails here, because proverbs do not translate:

The opening The correct completion
“illi eedo bil-mayy” “mish mitl illi eedo bil-naar”
“al-ʾird bi-ʿein immo” “ghazaal”
“tajree al-riyaah” “bimaa laa tashtahee al-sufun”
“ya daakhil bein al-basale w ʾishritha” “ma binoubak illa reehitha”
“illi ma byaʿref al-saʾr” “byishweeh”

A model that invents endings which are logical but do not exist has learned the structure of Arabic without having lived in it.

Item three: the greeting and sign-off test.

A simple and very sharp item. Ask: “Write me a short message to a work colleague apologising that I will not be able to attend tomorrow’s meeting.”

Watch four things:

Element Good sign Bad sign
Opening “marhaba” or “sabaah al-khair” “ʿazeezee al-zameel al-muhtaram”
Apology “baʿtizer minnak” “awadd an uʿrib ʿan iʿtidhaaree”
Reason Direct and short A full paragraph of justification
Closing “baʿtemed ʿaleik” or “minshoufak” “wa tafaddaloo bi-qubool faaʾiq al-ihtiraam”

The right-hand column is not a language error. It is a register error. A WhatsApp message to a colleague is not an official letter to a ministry. A model that cannot tell them apart learned Arabic from official and translated documents, which is the most abundant kind online.

Item four: measuring MSA leakage across six turns.

Start a conversation in Levantine and keep it going for six turns, recording the turn at which the model flips to Modern Standard Arabic:

Turn 1: “marhaba, keefak?”
Turn 2: “biddi asʾalak ʿan shaghle bil-shughl.”
Turn 3: “mudeeri talab minni taqreer w ana ma fhimt shu biddo bil-zabt.”
Turn 4: “laʾ, huwwe ma haka li. bas ʾaal jahhezli shi murattab.”
Turn 5: “tayyeb shu raʾyak aʿmal?”
Turn 6: “tamaam, bas khalleeha ʾaseere laʾenno ma ʿindi waʾt.”

The pattern I have watched repeat: turns one and two in Levantine, the third carrying formal words, the fourth fully formal. Record the turn number. It is a very useful and very fast metric, and I use it across this entire series.

Item five: the internal bilingualism test.

A bilingual model may think in one language and answer in another. Ask for a problem requiring steps, and ask for the steps in Arabic:

“I have a shop. I sold 17 items at 23 dinars each and paid 40 dinars for delivery. What is my net income? Write the steps in Arabic.”

Watch: are the steps genuinely Arabic, or numbers and English words? Did it keep “dinar” or convert it? Did it order the digits correctly? A model that gives you the right answer with half-English steps is thinking in English and translating.

Item six: the proper noun test.

Ask for a paragraph about an Arab figure who is not globally famous, such as the poet al-Mutanabbi, or Ibn al-Nafis, or Ghassan Kanafani. Then ask for the same paragraph in English. Compare the number of correct facts in each. If the English version is richer, the model stores its knowledge of your culture in somebody else’s language.


What Jais teaches us against what came before

Criterion Before Jais After it
Best open Arabic model BLOOMz at 42.9% Jais-chat at 48.4%
An open Arabic option at serious scale No Yes
Transparency about Arabic training data None A detailed table
Buildability for regional teams Difficult Direct
Evaluation on native Arabic benchmarks No Arrived later with ArabicMMLU

The third row is what I value most. Table 3 of the Jais paper breaks it down: original, translated, upsampled. That is transparency I have not found in most of what I read afterwards, including work two years newer.

And I will close with an observation I think matters for the region: transparency about weakness is more useful than an announcement about strength. A team that writes “a quarter of our Arabic is translated” gives the researchers who follow them a foundation to build on. A team that publishes one shiny number gives them a press headline.


The practical verdict

If you are choosing between a dedicated Arabic model and a global one, the rule I have arrived at: the dedicated Arabic model is better at register and cultural context, and the large global model is better at knowledge and reasoning. Choose by your task, not by the name.

If you are building Arabic training data, do not fill the gap with machine translation. Five billion tokens of native Arabic are worth more than twenty billion half of which is translated. And that is not only my opinion. The AraT5 paper proved it with numbers in 2022.

If you are reading an announcement about an Arabic model, ask for three numbers: how much original Arabic, how much translated, and how many times it was repeated. The announced figure alone tells you nothing.


Next in the series

One week after Jais, the UAE also announced Falcon 180B, the largest open model in the world at that point, trained on three and a half trillion tokens. The next article goes looking for Arabic in the model card. And does not find it.


Sources

  • Sengupta et al., Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models, arXiv:2308.16149, August 30, 2023
  • Cerebras, G42 and MBZUAI press release, August 2023: cerebras.ai
  • Koto et al., ArabicMMLU, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334
  • Huang et al., AceGPT, NAACL 2024: aclanthology.org/2024.naacl-long.450