The Letter Tax: Why ChatGPT Charged You Three Times More From Day One

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Model released: November 30, 2022 | Published: December 14, 2022

The numbers first

Metric Result What it measures Source
Cost of encoding Arabic against English 3 times Tokens needed for the same text carrying the same meaning Petrov, La Malfa, Torr and Bibi, NeurIPS 2023
Cost of Italian against English 1.6 times For comparison, a close European language Same paper
Cost of Bulgarian against English 2.6 times For comparison, a Cyrillic language Same paper
Widest gap between two languages in the study Up to 15 times How far the unfairness stretches across world languages Same paper
Tokenizer studied cl100k_base The tokenizer used by ChatGPT and GPT-4 Same paper
Reference corpus FLORES-200 Parallel text, human translated Same paper
ChatGPT release date November 30, 2022 The first mass encounter between Arabic speakers and a language model OpenAI launch page

The first number is the whole article. Everything after it is explanation.


What “three times” means in practice

Tokens needed for the same meaning, relative scale
English
1.0
Italian
1.6
Bulgarian
2.6
Arabic
3.0

And this is what that ratio turns into in your day:

Effect Detail
Cost The bill is counted in tokens. The same meaning in Arabic costs you three times as much
Speed The model generates token by token. An Arabic answer of equal meaning appears roughly three times slower
Context window A four thousand token window holds about a third of what it holds in English
Memory in a long chat The conversation history fills faster, so the model forgets the opening sooner
Output ceiling A long Arabic answer gets cut off before it finishes

Do this arithmetic yourself. It is simple division from the ratio above, not a new measurement: with an eight thousand token context you are putting roughly eight thousand units of meaning into it in English, and about two thousand six hundred in Arabic. You and your American colleague pay the same subscription and receive two different windows.


Card for the moment

Item Detail
Product ChatGPT
Release date November 30, 2022
Base From the GPT-3.5 series, which finished training in early 2022
Tokenizer cl100k_base
What was published about Arabic at the time Nothing. No number, no benchmark, no mention
How we learned the number Independent academic research, six months later

Look at that second-to-last line. It is not an aside, it is the rule across this entire series: Arabic numbers always come from outside the company that built the model, and they come late.


The ARAB-LENS reading: the unfairness starts before the intelligence

What bothers me most about the Arabic model debate is that it always starts from the wrong question: is the model smart in Arabic? The truth is that the unfairness begins a full step before intelligence enters the room.

A tokenizer is a small and relatively stupid piece of software whose job is to cut text into units. And it learns from the data it was fitted on. If that data is mostly English it will learn that “ing” and “tion” are whole units deserving one token each, and it will not learn that the Arabic definite article, the plural ending, or the attached pronoun deserve the same treatment. So it shatters the Arabic word into fragments.

And the consequence is not only an efficiency question, which is the part people miss. When a word is shattered into fragments that mean nothing on their own, the model has to reassemble the meaning from those fragments every single time. That is extra work, and it is paid for in accuracy, not only in money. My own view is that part of these models’ weakness in Arabic morphology traces back to exactly here, not to data scarcity alone.

Add to that the nature of Arabic itself. Arabic is a derivational, stacking language: a single word such as “wa-bi-madaarisihim” (and in their schools) carries a conjunction, a preposition, a definite article, a root, a pattern, a plural and a possessive pronoun. Seven units of meaning inside one word written with nine letters. A tokenizer built on English has no way to see that structure, so it treats it as an odd string of characters.

The last observation, and the harshest one: the paper reports that the gap between some language pairs reaches fifteen times. Arabic sits at three. Which means Arabic, for all its complaining, is relatively lucky in that ranking. Consider what the others are living with.


From the test notebook: measuring the letter tax on any model

You do not need a lab. You need two sentences and a ruler.

Item one: the parallel sentence.

Write the same meaning in Arabic and English and count the tokens for each with the model’s own tokenizer:

Arabic English
الاجتماع بكرا الساعة عشرة الصبح بمكتب المدير The meeting is tomorrow at ten in the morning in the manager’s office
لسا ما وصلني الملف، بتقدر تبعتلي إياه تاني؟ I have not received the file yet, could you send it to me again?
بدنا نعيد ترتيب الجدول قبل نهاية الأسبوع We need to rearrange the schedule before the end of the week

Record the ratio for each line. Then repeat with a passage of formal Arabic carrying full diacritics. You will find that vocalised Modern Standard Arabic is more expensive than dialect written plainly, because the diacritics add tokens without adding meaning.

Item two: the Arabizi tax.

Many people write Arabic in Latin letters and digits. Try the same text both ways:

In Arabic script In Arabizi The symbol and what it stands for
شو عم تعمل؟ shu 3am ta3mel? 3 stands for ʿayn
حبيبي كيفك؟ 7abibi kifak? 7 stands for ḥaaʾ
طيب خلص 6ayeb 5alas 6 for ṭaaʾ, 5 for khaaʾ
صار عندي موعد 9ar 3endi maw3ed 9 stands for ṣaad

The result I have seen repeatedly: Arabizi is often cheaper in tokens than the same sentence written in Arabic script. That is a cruel irony. The system pays you to abandon your alphabet.

Item three: the truncated context test.

Take a long Arabic passage that fills eighty percent of the context window, then ask a question about its opening line. Repeat with an English translation of the same passage. A model that answers correctly in English and fails in Arabic did not fail at comprehension. It failed at fitting.

Item four: the single word that equals a sentence.

Arabic agglutinates. A tokenizer built on English does not know where to cut. Try these words and count their tokens:

The word What it means in English Units of meaning inside it
a-fa-naqraʾuhaa and shall we read it 5: interrogative, conjunction, first person plural, verb, pronoun
wa-bi-madaarisihim and in their schools 6: conjunction, preposition, article, root, plural, pronoun
fa-sa-yaktuboonahaa so they will write it 5: conjunction, future marker, verb, plural subject, pronoun
li-yastakhrijoohaa for them to extract it 5
a-stasqeekumoohaa I ask you all for it to drink 6

Each of these is roughly six English words packed into one Arabic word. And yet when you count the tokens you will find the Arabic version is the more expensive one. That is not a linguistic paradox. It is the direct consequence of a tokenizer that never saw enough Arabic to learn that the “wa-bi-” prefix and the “-him” suffix are recurring units.

Item five: measuring the cost of diacritics.

Write the same line of verse three times: undiacriticised, partially diacriticised, fully diacriticised. Then count tokens.

Undiacriticised: al-ʿilm yarfaʿ baytan laa ʿimaad lah
Partially: al-ʿilmu yarfaʿu baytan laa ʿimaada lah
Fully: al-ʿilmu yarfaʿu baytan laa ʿimaada lahu, every vowel marked

The expected result: the third version can cost twice the first or more, and it carries exactly the same meaning to an Arabic reader. So if you are building an educational product that needs diacritics, know that you are paying twice: once for the letter tax and once for the vowel tax.


What has changed since that day, and what has not

I am writing this table because I want you to come back to it in two years and compare:

Item Its state in December 2022
Cost of an Arabic token Three times
An Arabic number published by the lab None
An adopted native Arabic benchmark None
Declared dialect support None
Open Arabic models at serious scale None
Number of Arabic speakers About half a billion

Five “none”s in front of half a billion people. I do not write that as a complaint, I write it as documentation. Because the first step in fixing anything is knowing where you started.

And the observation I want to leave with you: the letter tax is not a conspiracy, it is neglect. Nobody sat down and decided to make Arabic more expensive. Nobody simply sat down to make it cheaper. The tokenizer learned from the data that was available, the available data was English, and the result came out the way it came out. Which in my view is worse than a conspiracy, because a conspiracy can be exposed, whereas neglect first requires somebody to notice it.


The practical verdict

If you are building a product: budget in tokens, not words, and triple your Arabic estimate before you price anything. Then test the context window with real Arabic, not with a machine translation of English text, because machine translation produces Arabic that is shorter and simpler than the real thing.

If you are a user: know that the slowness of an Arabic answer is not your connection.

And if you are evaluating a model: do not compare answer lengths across the two languages directly. Compare density of meaning.


Next in the series

Four months after this moment GPT-4 arrived, and with it the first official Arabic number a major company ever published about its own model. The number was 80%, and it looked excellent. The next article explains why it is weaker than it looks, and what was hidden inside the way it was measured.


Sources

  • OpenAI, Introducing ChatGPT, November 30, 2022: openai.com/index/chatgpt
  • Petrov, La Malfa, Torr and Bibi, Language Model Tokenizers Introduce Unfairness Between Languages, NeurIPS 2023: arxiv.org/abs/2305.15425
  • NeurIPS 2023 Proceedings, volume 36: proceedings.neurips.cc