Before the Large Models Arrived: What Arabic Actually Had

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Work published: May 2022 (ACL 2022) | Published: June 5, 2022

The numbers first

Metric Result What it measures Source
AraBERT, training data 70 million sentences, about 24 GB The size of the Arabic input to the first serious Arabic model Antoun, Baly and Hajj, OSACT4 workshop, May 2020
AraBERT on HARD sentiment 96.2% against 95.7% for mBERT What a dedicated Arabic model buys you over a multilingual one Same paper
AraBERT on ANERcorp named entities 84.2 against 78.4 The gap widens as the task gets harder Same paper
ARBERT, Modern Standard Arabic data 61 GB, 6.5 billion tokens The largest clean MSA input of that generation Abdul-Mageed, Elmadany and Nagoudi, ACL-IJCNLP 2021
MARBERT, dialect data 1 billion Arabic tweets, 128 GB, 15.6 billion tokens The first serious attempt to teach a model the dialects Same paper
MARBERT-v2 on the ARLUE board 77.40 Average across six Arabic task clusters Same paper
AraT5, training data 248 GB, 29 billion tokens About 49% less than the 57 billion Arabic tokens mT5 saw Nagoudi, Elmadany and Abdul-Mageed, ACL 2022
AraT5 against mT5 on ARGEN Better on 52 of 59 test sets Arabic generation, not just classification Same paper
AraT5 on ARLUE 77.52 against 76.53 for MARBERT and 75.05 for mT5 The highest number of that generation Same paper

These are not chat model numbers. They belong to an entire earlier era, and if you do not know them you cannot understand where today’s problem came from.


What existed, and at what size

Arabic training data, in gigabytes
AraBERT 2020
24
ARBERT, MSA, 2021
61
AraT5, MSA, 2022
70
MARBERT, tweets, 2021
128
AraT5, combined, 2022
248

Look at the distance between the first line and the last: ten times in two years. By any normal measure that is healthy growth. The problem is that it happened in a parallel universe, while English was moving from billions of tokens to trillions.


Card for the era

Item Detail
Period 2020 to 2022
Institutions American University of Beirut, University of British Columbia
Architecture BERT for understanding, T5 for generation
Size Hundreds of millions of parameters, not billions
Strength Built for Arabic from scratch, not translated and not bolted on
Weakness No conversation, no instruction following, no reasoning
Real legacy The benchmarks and the data, not the models

The ARAB-LENS reading: the lesson we paid for twice

Something important happened in this era, and then it was completely forgotten.

The Canadian team at British Columbia did not build one model. They deliberately built two: ARBERT on 61 GB of clean Modern Standard Arabic, and MARBERT on a billion tweets. Why would a team go to that trouble twice? Because they understood early that Modern Standard Arabic and the dialects are not two registers of the same input. They are two different inputs, and each one needs its own training.

That understanding existed in 2021. And it was lost.

I mean lost literally. Every large Arabic model the region has built since, and you will see their numbers in the articles that follow, went back to mixing: one undifferentiated block of Arabic text, mostly formal, with whatever dialect happened to fall in. Then everyone is surprised that the model can read a newspaper column and cannot read a WhatsApp message.

But notice the detail that also exposes the limits of that era. How did the team collect its dialect data? With one filter: any tweet containing at least three Arabic words. That is the whole method. No dialect identification, no regional balance, no check that the tweet was in a dialect at all. A billion tweets came in through that filter, and from them the model learned what we called dialect.

Here is the opinion I will not soften: a three-Arabic-words filter is not a dialect collection methodology, it is an Arabic text collection methodology. The difference between those two things is the entire difference. And the same mistake is still being repeated in 2026, in newer and far more expensive forms.

The third observation is the one that matters most in practice. AraT5 beat mT5 on Arabic generation using about 49% less data. The ACL 2022 paper states it plainly: 29 billion Arabic tokens against mT5’s 57 billion, and a higher score. That is an early, clean proof that carefully chosen Arabic beats more Arabic scraped at random. Then the era of large models arrived and decided that size alone would do.


From the test notebook: how to tell whether an older model really reads a dialect

These items came out of the logic of that era, and I still run them on current models. Write them exactly as they are and use them.

Item one: the word that flips.

Take a single word whose grammatical function changes between Modern Standard Arabic and Levantine:

Word In Modern Standard Arabic In Levantine What it exposes
tayyeb, in “tayyeb, shu sar?” An adjective meaning good or tasty A discourse marker meaning all right, so If the model calls it an adjective it is reading in MSA
heyk Does not exist Like this, thus A word with no Standard Arabic root
3am, in “3am baktob” Paternal uncle A progressive marker The single most dangerous trap in Levantine
lissa Does not exist Still, not yet An entire tense marker

The test: write “3am baktob la3ammi” (I am writing to my uncle) and ask for a morphological analysis. A model that treats the first “3am” and the second as the same word has not learned the dialect. It has learned the shape of the letters.

Item two: qaf in four realizations.

Realization Rough region Example
qaal, with qaf MSA and some regions qaal li
ʾaal, with hamza Damascus, Beirut, Cairo ʾaal li
gaal, with hard g Gulf, Najd, Levantine bedouin gaal li
kaal, with kaf Parts of Palestine and rural areas kaal li

Put all four in one passage and ask the model to assign each sentence a region. The older models failed quietly and answered “Arabic”. The newer models answer confidently and are often wrong, which is worse.

Item three: the deletion test.

Take a dialect passage and delete every word that is not used in Modern Standard Arabic. If the passage is still coherent, it was formal Arabic wearing a dialect coat. Run it on these two:

First passage: “ana raye7 3al beit hallaʾ w ba3dein bshoufak.”
After deletion: “ana raye7 3al beit w ba3dein bshoufak.” Still alive. The real dialect content is low.

Second passage: “lissa 3am bestanna, w iza ma ija brou7 la7ali.”
After deletion: “w iza ma la7ali.” It collapsed. That is a real dialect.

This test is mine, and I use it as a fixed ruler across every article in this series.

Item four: ta marbuta in pausal form.

This one sorts models quickly, because it requires knowing how Arabic sounds and not only how it is spelled. The letter ta marbuta is pronounced as an h when you stop on it, and as a t when the word is joined to what follows.

Sentence Correct realization Common error
“shift madrase.” madraseh, an h, because we stopped madraset
“shift madraset al-7ayy.” madraset, a t, because we joined madraseh al-7ayy
“hayy sayyaara.” sayyaarah sayyaarat
“hayy sayyaarat abooy.” sayyaarat abooy sayyaarah abooy

Ask the model to write these four in Arabizi, that is in Latin letters. A model that writes madraseh in the first and madraset in the second understands the phonological structure. A model that writes the same form twice is reading the letter, not hearing it.

Item five: Amman and Oman.

Two words written with almost the same letters and different in everything else.

Word Pronunciation What it is Nationality
ʿAmmaan fatha on the ʿayn, doubled m The capital of Jordan Jordanian
ʿUmaan damma on the ʿayn, no doubling The Sultanate of Oman Omani

Write: “sa7bi min ʿaman bas ma baʿref iza ʿAmmaan walla ʿUmaan.” Then ask the model how it tells them apart. A correct answer names the gemination, the vowel and the nationality. A weak one apologises or guesses. I have run this item against every generation of model, and the success rate on it climbs far more slowly than on anything else.


What was actually lost between 2022 and 2023

I want to put this in one table, because it is the point that connects this article to everything that follows:

Property The AraBERT and MARBERT generation The first large-model generation
Separation of MSA from dialect Present and deliberate Gone
Arabic share of the data 100% Fractions of a percent
Curation quality Manual and reviewed Automated crawl
Benchmarks Native Arabic Machine-translated
Conversational ability None Excellent
Reasoning ability None Good
General knowledge Very limited Broad

Look at the two columns carefully. The older generation was better on the first four rows and the newer one is better on the last three. And because the last three are what a user sees, the new generation won and the old page was turned.

That, in my view, is the largest strategic error in the history of Arabic language processing: we did not have to choose. The curation and separation methodology of the first generation could have been carried inside the scale of the second. Nobody did it. Which is why in 2026 we are still rediscovering that Modern Standard Arabic and the dialects are two inputs and not one.


The practical verdict

If you are building an Arabic product today, three things from this era are still worth having:

First, ARLUE, ARGEN, ANERcorp and HARD are still available and still free, and they are native Arabic, not machine-translated. Use them for evaluation even if your model is a 2026 one. Many of the newer and more famous Arabic benchmarks are translated from English, and a later article shows what that cost.

Second, if your task is classification, named entity extraction or sentiment, a model the size of MARBERT may be all you need and will save you ninety percent of the cost. Do not call a hundred billion parameter model to tell you that a review is negative.

Third, the methodological lesson: separate Modern Standard Arabic from dialect in your data and in your evaluation. If your evaluation is a single number, you do not know where your product actually stands.


Next in the series

In November 2022 ChatGPT arrived, and hundreds of millions of Arabic speakers met a language model for the first time. The next article opens with a single number, but it is the number that explains why that first encounter was more expensive and slower for an Arabic speaker than for anyone else: three times.


Sources

  • Antoun, Baly and Hajj, AraBERT: Transformer-based Model for Arabic Language Understanding, OSACT4 workshop, May 2020: aclanthology.org/2020.osact-1.2
  • Abdul-Mageed, Elmadany and Nagoudi, ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic, ACL-IJCNLP 2021, August 2021: aclanthology.org/2021.acl-long.551
  • Nagoudi, Elmadany and Abdul-Mageed, AraT5: Text-to-Text Transformers for Arabic Language Generation, ACL 2022, May 2022: aclanthology.org/2022.acl-long.47