From 53 Tokens to 26: The Day the Letter Tax Was Cut

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Model released: May 13, 2024 | Published: May 27, 2024

The numbers first

Metric Result What it measures Source
Arabic, before the new tokenizer 53 tokens The same sample sentence GPT-4o launch page, May 13, 2024
Arabic, after the new tokenizer 26 tokens A reduction by a factor of two Same source
Declared improvement factor for Arabic 2.0x As OpenAI wrote it Same source
English, before and after 27 to 24 tokens, 1.1x The reference Same source
Gujarati, the largest gain 145 to 33, 4.4x Who benefited most Same source
Hindi 90 to 31, 2.9x For regional comparison Same source
Urdu 82 to 33, 2.5x A language in Arabic script Same source
Persian 61 to 32, 1.9x Another language in Arabic script Same source
The new tokenizer o200k_base Its technical name Same source

A necessary methodological note: these figures are for one sample sentence per language, not a corpus-level measurement. So do not say “Arabic became twice as cheap” as a general rule. Say “on the sentence OpenAI published, it dropped from 53 to 26.”


Before and after, on the published sentence

Token count for the sample sentence
Arabic, old tokenizer
53
Arabic, new tokenizer
26
English, old
27
English, new
24

And this is the figure worth computing: on this specific sentence, Arabic used to cost 1.96 times English and now costs 1.08 times. The gap has not been eliminated, but it has narrowed sharply.


Model card

Item Detail
Model GPT-4o
Lab OpenAI
Release date May 13, 2024
What is new Natively multimodal, faster, and a new tokenizer
Tokenizer o200k_base replacing cl100k_base
What was published about Arabic The tokenization table, an explicit official Arabic number
What was not published at launch Any Arabic performance benchmark for the May snapshot

The ARAB-LENS reading: the first time Arabic was counted

In article twenty of this series I wrote that the letter tax is not a conspiracy but neglect, and that neglect first requires somebody to notice it. This article is about the day somebody noticed.

Three observations, the first genuinely positive.

First: Arabic was in the table, by name, on day one.

That is not a cosmetic detail. The GPT-4o launch page displayed twenty languages with before and after counts, and Arabic was one of them, with a real Arabic sentence printed on the page. This is the first time in this entire series that I read an Arabic figure published by a major lab about its own model on launch day, without waiting two years for an academic paper.

Compare it with what came before. GPT-4 gave us a number inside a chart. Claude 3 did not break Arabic out. Falcon did not mention it. Mixtral left it off the list. And here: a clear text table, two numbers, and an improvement factor.

Second: who benefited most reveals who was treated worst.

Look at the ranking of improvements. Gujarati 4.4x, Telugu 3.5x, Tamil 3.3x, Marathi and Hindi 2.9x, Urdu 2.5x, then Arabic 2.0x. The languages that improved most are the ones that were worst off.

That confirms what I said in article twenty: Arabic, for all its complaining, was not at the bottom. Urdu, Hindi and Tamil were far worse. The new tokenizer was not built for Arabic. It was built largely for the languages of the subcontinent, and Arabic benefited on the way past.

Third, and it must be said: better tokenization is not better understanding.

This is the point I am afraid will get lost. Better tokenization means lower cost, higher speed, a wider effective context. It does not mean the model now knows Levantine, or understands that khalleena nshoof means no, or handles non-human plural agreement.

In fact the new tokenizer triggered a wide technical discussion, because some of the tokens it learned in certain languages came from low quality text. The lesson being that expanding a vocabulary without cleaning the data introduces problems of a different kind.


Where we stand: three moments compared

Moment Date The state of Arabic
ChatGPT November 2022 High letter tax, no published Arabic figure
GPT-4 March 2023 One number inside a chart, on a translated exam
GPT-4o May 2024 An explicit tokenization table, and no Arabic performance benchmark at launch

The direction is right but slow. Even in 2024, the question “how much Arabic does this model know?” still has no official answer on launch day.

A note about a later number you may encounter: there is today a published figure for GPT-4o on a version of MMLU translated into fourteen languages by professional human translators, Arabic among them. But that figure belongs to the November 2024 snapshot of the model, not the May release. Attaching it to the May launch is a common error. Which leads me to a general rule: a model’s name is not its identity. The snapshot is its identity.


From the test notebook: measuring the tokenizer’s effect on your own product

The published numbers are for one sentence. Your product is not one sentence. Here is a real measurement.

Step one: take a hundred texts from your product.

Not generic text. Your customers’ messages, your product descriptions, your FAQ. A hundred texts is enough for a stable measurement.

Step two: sort them into five categories.

Category Example Why it differs
Formal Arabic Terms and conditions Long sentences, standard vocabulary
Journalistic Arabic A product description Middle of the range
Written dialect A customer message Words outside the dictionary
Arabizi A message from a young customer Latin letters and digits
Mixed Arabic with English inside Code switching mid sentence

Count tokens per category and divide by word count. You will find that the third and fourth categories are far more expensive, because the tokenizer never learned them. And that tells you where your bill actually goes.

Step three: measure the difference between two tokenizers.

Run the hundred texts through two different tokenizers. The difference you find matters more than any launch figure, because it is the difference on your own data.

Step four: recompute your context budget.

If your application passes conversation history, work out how many turns actually fit in Arabic against English. Then set your truncation limit from the Arabic number, not the English one. The most common engineering error I have seen in Arabic applications: a context limit tuned on an English estimate, so the model starts forgetting the opening of the conversation at turn four instead of turn ten.

Final item: the Arabizi test again.

In article twenty we found that Arabizi is often cheaper than Arabic script. Rerun that test with the new tokenizer:

In Arabic script In Arabizi
“shu akhbaarak? keef el-shoghol?” “shu akhbarak? kif el shoghol?”
“biddi ehjez mawʿad bukra el-sobh” “biddi ehjez maw3ad bukra el sob7”
“el-talabiyye wislat naaʾsa” “el talabiyye wislat na2sa”

If the result flips and Arabic script becomes the cheaper one, that is the strongest signal that the tokenizer genuinely understands Arabic now, rather than that the numbers merely improved.

Item five: the real cost impact test.

Do not stop at counting tokens, compute the bill:

The item The method
Average prompt tokens From a hundred real samples
Average response tokens From a hundred real responses
Expected monthly requests From your own estimate
Difference between two tokenizers Multiply and subtract

Many Arab teams discovered after a tokenizer change that their budget had fallen by roughly a third. Few of them repriced to reflect it.

Item six: test whether quality improved, not only cost.

A hypothesis I recommend you test yourself: when tokenization improves, does grammar improve too?

The method: run ten sentences, each containing a long compound Arabic word such as wa-bi-madaarisihim, fa-sa-yaktuboonahaa, li-yastakhrijoohaa. Ask for a morphological analysis of each. Compare across two models with different tokenizers.

If the analysis improves with the better tokenizer, that supports the idea I raised in article twenty: shattering a word into fragments costs accuracy, not only money. If it does not improve, the two axes are entirely separate. I have no published figure that settles this, and I raise it as an open question rather than a conclusion.


The letter tax across three years

The moment What we know about the Arabic cost
ChatGPT, 2022 About 3 times English, NeurIPS 2023 research, measured on a corpus
GPT-4, 2023 Same tokenizer, so the same cost
GPT-4o, 2024 On the published sentence: 1.08 times English

An explicit methodological warning: you may not divide three by one point zero eight and declare a threefold improvement. The first number is a corpus measurement with the old tokenizer, the second is a count on one sentence with the new one. These are two measurements of different methods and one does not divide into the other.

I put this warning here because I see that arithmetic performed often in Arabic coverage, producing numbers with no basis.


The practical verdict

Recompute your cost with every tokenizer change. Old budget estimates become wrong by a factor of two, in both directions.

And do not confuse cheap tokenization with language quality. These are two entirely separate axes, and a model that is cheap in Arabic can remain ignorant of it.

Always name the snapshot. “GPT-4o” on its own is a family name, not a model. And any comparison without a snapshot date is a comparison without meaning.


Next in the series

On roughly the same day, the Technology Innovation Institute in Abu Dhabi released Falcon 2 at eleven billion parameters. Its technical report contains exactly one Arabic number, and it is a number that deserves to be read slowly: 25.32. Random guessing on that test is 25.


Sources

  • OpenAI, Hello GPT-4o, May 13, 2024: openai.com/index/hello-gpt-4o
  • The tokenization comparison table on the same page
  • openai/simple-evals, multilingual MMMLU results, for the November 2024 snapshot reference: github.com/openai/simple-evals