The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Cost of encoding Arabic against English | 3 times | Tokens needed for the same text carrying the same meaning | Petrov, La Malfa, Torr and Bibi, NeurIPS 2023 |
| Cost of Italian against English | 1.6 times | For comparison, a close European language | Same paper |
| Cost of Bulgarian against English | 2.6 times | For comparison, a Cyrillic language | Same paper |
| Widest gap between two languages in the study | Up to 15 times | How far the unfairness stretches across world languages | Same paper |
| Tokenizer studied | cl100k_base | The tokenizer used by ChatGPT and GPT-4 | Same paper |
| Reference corpus | FLORES-200 | Parallel text, human translated | Same paper |
| ChatGPT release date | November 30, 2022 | The first mass encounter between Arabic speakers and a language model | OpenAI launch page |
The first number is the whole article. Everything after it is explanation.
What “three times” means in practice
And this is what that ratio turns into in your day:
| Effect | Detail |
|---|---|
| Cost | The bill is counted in tokens. The same meaning in Arabic costs you three times as much |
| Speed | The model generates token by token. An Arabic answer of equal meaning appears roughly three times slower |
| Context window | A four thousand token window holds about a third of what it holds in English |
| Memory in a long chat | The conversation history fills faster, so the model forgets the opening sooner |
| Output ceiling | A long Arabic answer gets cut off before it finishes |
Do this arithmetic yourself. It is simple division from the ratio above, not a new measurement: with an eight thousand token context you are putting roughly eight thousand units of meaning into it in English, and about two thousand six hundred in Arabic. You and your American colleague pay the same subscription and receive two different windows.
Card for the moment
| Item | Detail |
|---|---|
| Product | ChatGPT |
| Release date | November 30, 2022 |
| Base | From the GPT-3.5 series, which finished training in early 2022 |
| Tokenizer | cl100k_base |
| What was published about Arabic at the time | Nothing. No number, no benchmark, no mention |
| How we learned the number | Independent academic research, six months later |
Look at that second-to-last line. It is not an aside, it is the rule across this entire series: Arabic numbers always come from outside the company that built the model, and they come late.
The ARAB-LENS reading: the unfairness starts before the intelligence
What bothers me most about the Arabic model debate is that it always starts from the wrong question: is the model smart in Arabic? The truth is that the unfairness begins a full step before intelligence enters the room.
A tokenizer is a small and relatively stupid piece of software whose job is to cut text into units. And it learns from the data it was fitted on. If that data is mostly English it will learn that “ing” and “tion” are whole units deserving one token each, and it will not learn that the Arabic definite article, the plural ending, or the attached pronoun deserve the same treatment. So it shatters the Arabic word into fragments.
And the consequence is not only an efficiency question, which is the part people miss. When a word is shattered into fragments that mean nothing on their own, the model has to reassemble the meaning from those fragments every single time. That is extra work, and it is paid for in accuracy, not only in money. My own view is that part of these models’ weakness in Arabic morphology traces back to exactly here, not to data scarcity alone.
Add to that the nature of Arabic itself. Arabic is a derivational, stacking language: a single word such as “wa-bi-madaarisihim” (and in their schools) carries a conjunction, a preposition, a definite article, a root, a pattern, a plural and a possessive pronoun. Seven units of meaning inside one word written with nine letters. A tokenizer built on English has no way to see that structure, so it treats it as an odd string of characters.
The last observation, and the harshest one: the paper reports that the gap between some language pairs reaches fifteen times. Arabic sits at three. Which means Arabic, for all its complaining, is relatively lucky in that ranking. Consider what the others are living with.
From the test notebook: measuring the letter tax on any model
You do not need a lab. You need two sentences and a ruler.
Item one: the parallel sentence.
Write the same meaning in Arabic and English and count the tokens for each with the model’s own tokenizer:
| Arabic | English |
|---|---|
| الاجتماع بكرا الساعة عشرة الصبح بمكتب المدير | The meeting is tomorrow at ten in the morning in the manager’s office |
| لسا ما وصلني الملف، بتقدر تبعتلي إياه تاني؟ | I have not received the file yet, could you send it to me again? |
| بدنا نعيد ترتيب الجدول قبل نهاية الأسبوع | We need to rearrange the schedule before the end of the week |
Record the ratio for each line. Then repeat with a passage of formal Arabic carrying full diacritics. You will find that vocalised Modern Standard Arabic is more expensive than dialect written plainly, because the diacritics add tokens without adding meaning.
Item two: the Arabizi tax.
Many people write Arabic in Latin letters and digits. Try the same text both ways:
| In Arabic script | In Arabizi | The symbol and what it stands for |
|---|---|---|
| شو عم تعمل؟ | shu 3am ta3mel? | 3 stands for ʿayn |
| حبيبي كيفك؟ | 7abibi kifak? | 7 stands for ḥaaʾ |
| طيب خلص | 6ayeb 5alas | 6 for ṭaaʾ, 5 for khaaʾ |
| صار عندي موعد | 9ar 3endi maw3ed | 9 stands for ṣaad |
The result I have seen repeatedly: Arabizi is often cheaper in tokens than the same sentence written in Arabic script. That is a cruel irony. The system pays you to abandon your alphabet.
Item three: the truncated context test.
Take a long Arabic passage that fills eighty percent of the context window, then ask a question about its opening line. Repeat with an English translation of the same passage. A model that answers correctly in English and fails in Arabic did not fail at comprehension. It failed at fitting.
Item four: the single word that equals a sentence.
Arabic agglutinates. A tokenizer built on English does not know where to cut. Try these words and count their tokens:
| The word | What it means in English | Units of meaning inside it |
|---|---|---|
| a-fa-naqraʾuhaa | and shall we read it | 5: interrogative, conjunction, first person plural, verb, pronoun |
| wa-bi-madaarisihim | and in their schools | 6: conjunction, preposition, article, root, plural, pronoun |
| fa-sa-yaktuboonahaa | so they will write it | 5: conjunction, future marker, verb, plural subject, pronoun |
| li-yastakhrijoohaa | for them to extract it | 5 |
| a-stasqeekumoohaa | I ask you all for it to drink | 6 |
Each of these is roughly six English words packed into one Arabic word. And yet when you count the tokens you will find the Arabic version is the more expensive one. That is not a linguistic paradox. It is the direct consequence of a tokenizer that never saw enough Arabic to learn that the “wa-bi-” prefix and the “-him” suffix are recurring units.
Item five: measuring the cost of diacritics.
Write the same line of verse three times: undiacriticised, partially diacriticised, fully diacriticised. Then count tokens.
Undiacriticised: al-ʿilm yarfaʿ baytan laa ʿimaad lah
Partially: al-ʿilmu yarfaʿu baytan laa ʿimaada lah
Fully: al-ʿilmu yarfaʿu baytan laa ʿimaada lahu, every vowel marked
The expected result: the third version can cost twice the first or more, and it carries exactly the same meaning to an Arabic reader. So if you are building an educational product that needs diacritics, know that you are paying twice: once for the letter tax and once for the vowel tax.
What has changed since that day, and what has not
I am writing this table because I want you to come back to it in two years and compare:
| Item | Its state in December 2022 |
|---|---|
| Cost of an Arabic token | Three times |
| An Arabic number published by the lab | None |
| An adopted native Arabic benchmark | None |
| Declared dialect support | None |
| Open Arabic models at serious scale | None |
| Number of Arabic speakers | About half a billion |
Five “none”s in front of half a billion people. I do not write that as a complaint, I write it as documentation. Because the first step in fixing anything is knowing where you started.
And the observation I want to leave with you: the letter tax is not a conspiracy, it is neglect. Nobody sat down and decided to make Arabic more expensive. Nobody simply sat down to make it cheaper. The tokenizer learned from the data that was available, the available data was English, and the result came out the way it came out. Which in my view is worse than a conspiracy, because a conspiracy can be exposed, whereas neglect first requires somebody to notice it.
The practical verdict
If you are building a product: budget in tokens, not words, and triple your Arabic estimate before you price anything. Then test the context window with real Arabic, not with a machine translation of English text, because machine translation produces Arabic that is shorter and simpler than the real thing.
If you are a user: know that the slowness of an Arabic answer is not your connection.
And if you are evaluating a model: do not compare answer lengths across the two languages directly. Compare density of meaning.
Next in the series
Four months after this moment GPT-4 arrived, and with it the first official Arabic number a major company ever published about its own model. The number was 80%, and it looked excellent. The next article explains why it is weaker than it looks, and what was hidden inside the way it was measured.
Sources
- OpenAI, Introducing ChatGPT, November 30, 2022: openai.com/index/chatgpt
- Petrov, La Malfa, Torr and Bibi, Language Model Tokenizers Introduce Unfairness Between Languages, NeurIPS 2023: arxiv.org/abs/2305.15425
- NeurIPS 2023 Proceedings, volume 36: proceedings.neurips.cc