The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| AraBERT, training data | 70 million sentences, about 24 GB | The size of the Arabic input to the first serious Arabic model | Antoun, Baly and Hajj, OSACT4 workshop, May 2020 |
| AraBERT on HARD sentiment | 96.2% against 95.7% for mBERT | What a dedicated Arabic model buys you over a multilingual one | Same paper |
| AraBERT on ANERcorp named entities | 84.2 against 78.4 | The gap widens as the task gets harder | Same paper |
| ARBERT, Modern Standard Arabic data | 61 GB, 6.5 billion tokens | The largest clean MSA input of that generation | Abdul-Mageed, Elmadany and Nagoudi, ACL-IJCNLP 2021 |
| MARBERT, dialect data | 1 billion Arabic tweets, 128 GB, 15.6 billion tokens | The first serious attempt to teach a model the dialects | Same paper |
| MARBERT-v2 on the ARLUE board | 77.40 | Average across six Arabic task clusters | Same paper |
| AraT5, training data | 248 GB, 29 billion tokens | About 49% less than the 57 billion Arabic tokens mT5 saw | Nagoudi, Elmadany and Abdul-Mageed, ACL 2022 |
| AraT5 against mT5 on ARGEN | Better on 52 of 59 test sets | Arabic generation, not just classification | Same paper |
| AraT5 on ARLUE | 77.52 against 76.53 for MARBERT and 75.05 for mT5 | The highest number of that generation | Same paper |
These are not chat model numbers. They belong to an entire earlier era, and if you do not know them you cannot understand where today’s problem came from.
What existed, and at what size
Look at the distance between the first line and the last: ten times in two years. By any normal measure that is healthy growth. The problem is that it happened in a parallel universe, while English was moving from billions of tokens to trillions.
Card for the era
| Item | Detail |
|---|---|
| Period | 2020 to 2022 |
| Institutions | American University of Beirut, University of British Columbia |
| Architecture | BERT for understanding, T5 for generation |
| Size | Hundreds of millions of parameters, not billions |
| Strength | Built for Arabic from scratch, not translated and not bolted on |
| Weakness | No conversation, no instruction following, no reasoning |
| Real legacy | The benchmarks and the data, not the models |
The ARAB-LENS reading: the lesson we paid for twice
Something important happened in this era, and then it was completely forgotten.
The Canadian team at British Columbia did not build one model. They deliberately built two: ARBERT on 61 GB of clean Modern Standard Arabic, and MARBERT on a billion tweets. Why would a team go to that trouble twice? Because they understood early that Modern Standard Arabic and the dialects are not two registers of the same input. They are two different inputs, and each one needs its own training.
That understanding existed in 2021. And it was lost.
I mean lost literally. Every large Arabic model the region has built since, and you will see their numbers in the articles that follow, went back to mixing: one undifferentiated block of Arabic text, mostly formal, with whatever dialect happened to fall in. Then everyone is surprised that the model can read a newspaper column and cannot read a WhatsApp message.
But notice the detail that also exposes the limits of that era. How did the team collect its dialect data? With one filter: any tweet containing at least three Arabic words. That is the whole method. No dialect identification, no regional balance, no check that the tweet was in a dialect at all. A billion tweets came in through that filter, and from them the model learned what we called dialect.
Here is the opinion I will not soften: a three-Arabic-words filter is not a dialect collection methodology, it is an Arabic text collection methodology. The difference between those two things is the entire difference. And the same mistake is still being repeated in 2026, in newer and far more expensive forms.
The third observation is the one that matters most in practice. AraT5 beat mT5 on Arabic generation using about 49% less data. The ACL 2022 paper states it plainly: 29 billion Arabic tokens against mT5’s 57 billion, and a higher score. That is an early, clean proof that carefully chosen Arabic beats more Arabic scraped at random. Then the era of large models arrived and decided that size alone would do.
From the test notebook: how to tell whether an older model really reads a dialect
These items came out of the logic of that era, and I still run them on current models. Write them exactly as they are and use them.
Item one: the word that flips.
Take a single word whose grammatical function changes between Modern Standard Arabic and Levantine:
| Word | In Modern Standard Arabic | In Levantine | What it exposes |
|---|---|---|---|
| tayyeb, in “tayyeb, shu sar?” | An adjective meaning good or tasty | A discourse marker meaning all right, so | If the model calls it an adjective it is reading in MSA |
| heyk | Does not exist | Like this, thus | A word with no Standard Arabic root |
| 3am, in “3am baktob” | Paternal uncle | A progressive marker | The single most dangerous trap in Levantine |
| lissa | Does not exist | Still, not yet | An entire tense marker |
The test: write “3am baktob la3ammi” (I am writing to my uncle) and ask for a morphological analysis. A model that treats the first “3am” and the second as the same word has not learned the dialect. It has learned the shape of the letters.
Item two: qaf in four realizations.
| Realization | Rough region | Example |
|---|---|---|
| qaal, with qaf | MSA and some regions | qaal li |
| ʾaal, with hamza | Damascus, Beirut, Cairo | ʾaal li |
| gaal, with hard g | Gulf, Najd, Levantine bedouin | gaal li |
| kaal, with kaf | Parts of Palestine and rural areas | kaal li |
Put all four in one passage and ask the model to assign each sentence a region. The older models failed quietly and answered “Arabic”. The newer models answer confidently and are often wrong, which is worse.
Item three: the deletion test.
Take a dialect passage and delete every word that is not used in Modern Standard Arabic. If the passage is still coherent, it was formal Arabic wearing a dialect coat. Run it on these two:
First passage: “ana raye7 3al beit hallaʾ w ba3dein bshoufak.”
After deletion: “ana raye7 3al beit w ba3dein bshoufak.” Still alive. The real dialect content is low.
Second passage: “lissa 3am bestanna, w iza ma ija brou7 la7ali.”
After deletion: “w iza ma la7ali.” It collapsed. That is a real dialect.
This test is mine, and I use it as a fixed ruler across every article in this series.
Item four: ta marbuta in pausal form.
This one sorts models quickly, because it requires knowing how Arabic sounds and not only how it is spelled. The letter ta marbuta is pronounced as an h when you stop on it, and as a t when the word is joined to what follows.
| Sentence | Correct realization | Common error |
|---|---|---|
| “shift madrase.” | madraseh, an h, because we stopped | madraset |
| “shift madraset al-7ayy.” | madraset, a t, because we joined | madraseh al-7ayy |
| “hayy sayyaara.” | sayyaarah | sayyaarat |
| “hayy sayyaarat abooy.” | sayyaarat abooy | sayyaarah abooy |
Ask the model to write these four in Arabizi, that is in Latin letters. A model that writes madraseh in the first and madraset in the second understands the phonological structure. A model that writes the same form twice is reading the letter, not hearing it.
Item five: Amman and Oman.
Two words written with almost the same letters and different in everything else.
| Word | Pronunciation | What it is | Nationality |
|---|---|---|---|
| ʿAmmaan | fatha on the ʿayn, doubled m | The capital of Jordan | Jordanian |
| ʿUmaan | damma on the ʿayn, no doubling | The Sultanate of Oman | Omani |
Write: “sa7bi min ʿaman bas ma baʿref iza ʿAmmaan walla ʿUmaan.” Then ask the model how it tells them apart. A correct answer names the gemination, the vowel and the nationality. A weak one apologises or guesses. I have run this item against every generation of model, and the success rate on it climbs far more slowly than on anything else.
What was actually lost between 2022 and 2023
I want to put this in one table, because it is the point that connects this article to everything that follows:
| Property | The AraBERT and MARBERT generation | The first large-model generation |
|---|---|---|
| Separation of MSA from dialect | Present and deliberate | Gone |
| Arabic share of the data | 100% | Fractions of a percent |
| Curation quality | Manual and reviewed | Automated crawl |
| Benchmarks | Native Arabic | Machine-translated |
| Conversational ability | None | Excellent |
| Reasoning ability | None | Good |
| General knowledge | Very limited | Broad |
Look at the two columns carefully. The older generation was better on the first four rows and the newer one is better on the last three. And because the last three are what a user sees, the new generation won and the old page was turned.
That, in my view, is the largest strategic error in the history of Arabic language processing: we did not have to choose. The curation and separation methodology of the first generation could have been carried inside the scale of the second. Nobody did it. Which is why in 2026 we are still rediscovering that Modern Standard Arabic and the dialects are two inputs and not one.
The practical verdict
If you are building an Arabic product today, three things from this era are still worth having:
First, ARLUE, ARGEN, ANERcorp and HARD are still available and still free, and they are native Arabic, not machine-translated. Use them for evaluation even if your model is a 2026 one. Many of the newer and more famous Arabic benchmarks are translated from English, and a later article shows what that cost.
Second, if your task is classification, named entity extraction or sentiment, a model the size of MARBERT may be all you need and will save you ninety percent of the cost. Do not call a hundred billion parameter model to tell you that a review is negative.
Third, the methodological lesson: separate Modern Standard Arabic from dialect in your data and in your evaluation. If your evaluation is a single number, you do not know where your product actually stands.
Next in the series
In November 2022 ChatGPT arrived, and hundreds of millions of Arabic speakers met a language model for the first time. The next article opens with a single number, but it is the number that explains why that first encounter was more expensive and slower for an Arabic speaker than for anyone else: three times.
Sources
- Antoun, Baly and Hajj, AraBERT: Transformer-based Model for Arabic Language Understanding, OSACT4 workshop, May 2020: aclanthology.org/2020.osact-1.2
- Abdul-Mageed, Elmadany and Nagoudi, ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic, ACL-IJCNLP 2021, August 2021: aclanthology.org/2021.acl-long.551
- Nagoudi, Elmadany and Abdul-Mageed, AraT5: Text-to-Text Transformers for Arabic Language Generation, ACL 2022, May 2022: aclanthology.org/2022.acl-long.47