The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Arabic, before the new tokenizer | 53 tokens | The same sample sentence | GPT-4o launch page, May 13, 2024 |
| Arabic, after the new tokenizer | 26 tokens | A reduction by a factor of two | Same source |
| Declared improvement factor for Arabic | 2.0x | As OpenAI wrote it | Same source |
| English, before and after | 27 to 24 tokens, 1.1x | The reference | Same source |
| Gujarati, the largest gain | 145 to 33, 4.4x | Who benefited most | Same source |
| Hindi | 90 to 31, 2.9x | For regional comparison | Same source |
| Urdu | 82 to 33, 2.5x | A language in Arabic script | Same source |
| Persian | 61 to 32, 1.9x | Another language in Arabic script | Same source |
| The new tokenizer | o200k_base | Its technical name | Same source |
A necessary methodological note: these figures are for one sample sentence per language, not a corpus-level measurement. So do not say “Arabic became twice as cheap” as a general rule. Say “on the sentence OpenAI published, it dropped from 53 to 26.”
Before and after, on the published sentence
And this is the figure worth computing: on this specific sentence, Arabic used to cost 1.96 times English and now costs 1.08 times. The gap has not been eliminated, but it has narrowed sharply.
Model card
| Item | Detail |
|---|---|
| Model | GPT-4o |
| Lab | OpenAI |
| Release date | May 13, 2024 |
| What is new | Natively multimodal, faster, and a new tokenizer |
| Tokenizer | o200k_base replacing cl100k_base |
| What was published about Arabic | The tokenization table, an explicit official Arabic number |
| What was not published at launch | Any Arabic performance benchmark for the May snapshot |
The ARAB-LENS reading: the first time Arabic was counted
In article twenty of this series I wrote that the letter tax is not a conspiracy but neglect, and that neglect first requires somebody to notice it. This article is about the day somebody noticed.
Three observations, the first genuinely positive.
First: Arabic was in the table, by name, on day one.
That is not a cosmetic detail. The GPT-4o launch page displayed twenty languages with before and after counts, and Arabic was one of them, with a real Arabic sentence printed on the page. This is the first time in this entire series that I read an Arabic figure published by a major lab about its own model on launch day, without waiting two years for an academic paper.
Compare it with what came before. GPT-4 gave us a number inside a chart. Claude 3 did not break Arabic out. Falcon did not mention it. Mixtral left it off the list. And here: a clear text table, two numbers, and an improvement factor.
Second: who benefited most reveals who was treated worst.
Look at the ranking of improvements. Gujarati 4.4x, Telugu 3.5x, Tamil 3.3x, Marathi and Hindi 2.9x, Urdu 2.5x, then Arabic 2.0x. The languages that improved most are the ones that were worst off.
That confirms what I said in article twenty: Arabic, for all its complaining, was not at the bottom. Urdu, Hindi and Tamil were far worse. The new tokenizer was not built for Arabic. It was built largely for the languages of the subcontinent, and Arabic benefited on the way past.
Third, and it must be said: better tokenization is not better understanding.
This is the point I am afraid will get lost. Better tokenization means lower cost, higher speed, a wider effective context. It does not mean the model now knows Levantine, or understands that khalleena nshoof means no, or handles non-human plural agreement.
In fact the new tokenizer triggered a wide technical discussion, because some of the tokens it learned in certain languages came from low quality text. The lesson being that expanding a vocabulary without cleaning the data introduces problems of a different kind.
Where we stand: three moments compared
| Moment | Date | The state of Arabic |
|---|---|---|
| ChatGPT | November 2022 | High letter tax, no published Arabic figure |
| GPT-4 | March 2023 | One number inside a chart, on a translated exam |
| GPT-4o | May 2024 | An explicit tokenization table, and no Arabic performance benchmark at launch |
The direction is right but slow. Even in 2024, the question “how much Arabic does this model know?” still has no official answer on launch day.
A note about a later number you may encounter: there is today a published figure for GPT-4o on a version of MMLU translated into fourteen languages by professional human translators, Arabic among them. But that figure belongs to the November 2024 snapshot of the model, not the May release. Attaching it to the May launch is a common error. Which leads me to a general rule: a model’s name is not its identity. The snapshot is its identity.
From the test notebook: measuring the tokenizer’s effect on your own product
The published numbers are for one sentence. Your product is not one sentence. Here is a real measurement.
Step one: take a hundred texts from your product.
Not generic text. Your customers’ messages, your product descriptions, your FAQ. A hundred texts is enough for a stable measurement.
Step two: sort them into five categories.
| Category | Example | Why it differs |
|---|---|---|
| Formal Arabic | Terms and conditions | Long sentences, standard vocabulary |
| Journalistic Arabic | A product description | Middle of the range |
| Written dialect | A customer message | Words outside the dictionary |
| Arabizi | A message from a young customer | Latin letters and digits |
| Mixed | Arabic with English inside | Code switching mid sentence |
Count tokens per category and divide by word count. You will find that the third and fourth categories are far more expensive, because the tokenizer never learned them. And that tells you where your bill actually goes.
Step three: measure the difference between two tokenizers.
Run the hundred texts through two different tokenizers. The difference you find matters more than any launch figure, because it is the difference on your own data.
Step four: recompute your context budget.
If your application passes conversation history, work out how many turns actually fit in Arabic against English. Then set your truncation limit from the Arabic number, not the English one. The most common engineering error I have seen in Arabic applications: a context limit tuned on an English estimate, so the model starts forgetting the opening of the conversation at turn four instead of turn ten.
Final item: the Arabizi test again.
In article twenty we found that Arabizi is often cheaper than Arabic script. Rerun that test with the new tokenizer:
| In Arabic script | In Arabizi |
|---|---|
| “shu akhbaarak? keef el-shoghol?” | “shu akhbarak? kif el shoghol?” |
| “biddi ehjez mawʿad bukra el-sobh” | “biddi ehjez maw3ad bukra el sob7” |
| “el-talabiyye wislat naaʾsa” | “el talabiyye wislat na2sa” |
If the result flips and Arabic script becomes the cheaper one, that is the strongest signal that the tokenizer genuinely understands Arabic now, rather than that the numbers merely improved.
Item five: the real cost impact test.
Do not stop at counting tokens, compute the bill:
| The item | The method |
|---|---|
| Average prompt tokens | From a hundred real samples |
| Average response tokens | From a hundred real responses |
| Expected monthly requests | From your own estimate |
| Difference between two tokenizers | Multiply and subtract |
Many Arab teams discovered after a tokenizer change that their budget had fallen by roughly a third. Few of them repriced to reflect it.
Item six: test whether quality improved, not only cost.
A hypothesis I recommend you test yourself: when tokenization improves, does grammar improve too?
The method: run ten sentences, each containing a long compound Arabic word such as wa-bi-madaarisihim, fa-sa-yaktuboonahaa, li-yastakhrijoohaa. Ask for a morphological analysis of each. Compare across two models with different tokenizers.
If the analysis improves with the better tokenizer, that supports the idea I raised in article twenty: shattering a word into fragments costs accuracy, not only money. If it does not improve, the two axes are entirely separate. I have no published figure that settles this, and I raise it as an open question rather than a conclusion.
The letter tax across three years
| The moment | What we know about the Arabic cost |
|---|---|
| ChatGPT, 2022 | About 3 times English, NeurIPS 2023 research, measured on a corpus |
| GPT-4, 2023 | Same tokenizer, so the same cost |
| GPT-4o, 2024 | On the published sentence: 1.08 times English |
An explicit methodological warning: you may not divide three by one point zero eight and declare a threefold improvement. The first number is a corpus measurement with the old tokenizer, the second is a count on one sentence with the new one. These are two measurements of different methods and one does not divide into the other.
I put this warning here because I see that arithmetic performed often in Arabic coverage, producing numbers with no basis.
The practical verdict
Recompute your cost with every tokenizer change. Old budget estimates become wrong by a factor of two, in both directions.
And do not confuse cheap tokenization with language quality. These are two entirely separate axes, and a model that is cheap in Arabic can remain ignorant of it.
Always name the snapshot. “GPT-4o” on its own is a family name, not a model. And any comparison without a snapshot date is a comparison without meaning.
Next in the series
On roughly the same day, the Technology Innovation Institute in Abu Dhabi released Falcon 2 at eleven billion parameters. Its technical report contains exactly one Arabic number, and it is a number that deserves to be read slowly: 25.32. Random guessing on that test is 25.
Sources
- OpenAI, Hello GPT-4o, May 13, 2024: openai.com/index/hello-gpt-4o
- The tokenization comparison table on the same page
- openai/simple-evals, multilingual MMMLU results, for the November 2024 snapshot reference: github.com/openai/simple-evals