75 Percent on an Arabic Leaderboard: What the Falcon-H1-Arabic Number Measures, and What It Hides

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: January 5, 2026 | Published: January 19, 2026

The numbers first

ModelSizeContextOALLSource
Falcon-H1-Arabic 34B34B256K tokens~75.36%Technology Innovation Institute, January 2026
Falcon-H1-Arabic 7B7B256K tokens71.47%Same source
Falcon-H1-Arabic 3B3B128K tokens~62%Same source

And the comparisons as the developer states them:

SizeAhead ofStated margin
3BGemma-4B and Qwen3-4BAbout ten points
7BFanar-9B and ALLaM-7BStated lead, no specific figure
34BLlama-3.3-70B and AceGPT2-32BStated lead, no specific figure
OALL: specialization versus scale
Falcon-H1-Arabic 34B
75.4
Falcon-H1-Arabic 7B
71.5
Falcon-H1-Arabic 3B
62.0
A seven billion parameter model passing models
ten times its size, when the measurement is Arabic

That table says one thing plainly: in Arabic, specialization beats scale. A seven billion parameter model that runs on a single machine outperforms models with tens of billions when the question is natively Arabic rather than translated. That is the most important structural result to come out of the region this past year.

Then comes the question this article exists for: seventy five percent of what?


What the OALL leaderboard actually measures

OALL version two is a serious step forward, and the people behind it deserve credit for one brave decision: they removed the machine-translated benchmarks. Their stated reason was direct, that many tasks were straight translations from English which introduced linguistic and contextual mismatches. They also dropped benchmarks that had saturated to the point of no longer separating models.

The leaderboard now carries native benchmarks: ArabicMMLU for school knowledge, MadinahQA for language and grammar, AraTrust for safety, ALRAGE for grounded question answering, alongside AlGhafa, EXAMS and Belebele.

Now look at that list through the eyes of someone building an Arabic product, and ask: where is dialect?

ARAB-LENS axisDoes OALL measure it?
Language: grammar, morphology, orthographyYes, via MadinahQA
General knowledge in ArabicYes, via ArabicMMLU and EXAMS
SafetyYes, via AraTrust
Grounded answeringYes, via ALRAGE
Dialect understandingNo
Dialect generationNo
PragmaticsNo
Multi-turn conversationNo
SpeechNo

Five of nine axes sit entirely outside the measurement. Which makes “this model scored 75 on Arabic” both true and incomplete: true because the number is real and its source is public, incomplete because the number knows nothing about half of what you need.


Model card

FieldDetail
NameFalcon-H1-Arabic
LabTechnology Innovation Institute, Abu Dhabi
ReleasedJanuary 5, 2026
Sizes3B, 7B, 34B
ArchitectureHybrid Mamba and Transformer, attention and SSM running in parallel inside each block, outputs concatenated
PredecessorFalcon-Arabic, released a few months earlier
Axis under testARAB-DIALECT: identification and discrimination

The hybrid architecture deserves a note. Running attention and a state space mechanism side by side inside one block buys long context at lower cost, and that matters specifically for Arabic, a derivational language with long words whose texts consume more tokens than English for the same meaning.


A five-level dialect test

Since the leaderboard does not measure dialect, here is what I measure. These items are written in Levantine and Jordanian Arabic, and anyone can rerun them against the same model.

Level 1, identification. “halla’ ma ili khili’ ahki maʿ hada.” Required: name the dialect family with a confidence level. My rule: no penalty for saying “Levantine” rather than “Syrian,” because a dialect is not a nationality. The penalty is for overconfidence.

Level 2, lexical. Does it know khili’ here means mood, not creation? Does it know hada means anyone?

Level 3, meaning in context. Put the word ʿammar in three contexts: a proper name, construction, and the blessing “bil-ʿammar ya rabb.” Three different answers or a fail.

Level 4, conversion. “madha tafʿal al-‘an?” into Syrian, then Jordanian, then Egyptian, then Moroccan, with the requirement that they differ structurally rather than lexically.

Level 5, register. The same meaning to a friend, a manager, an elder, on WhatsApp, and in a customer support reply.

I add a sixth item specific to Arabic-specialized models, because it is the trap they fall into: the dialect confidence trap. Give it a northern rural Jordanian sentence using gal rather than qal, and ask where the speaker is from. A model trained mostly on urban Levantine will answer “Syrian” with high confidence, and that is a particular kind of error: not ignorance of Arabic, but generalizing a training sample across a region far wider than the sample.


The comparison worth reading carefully

Here is a detail that needs precision. On HELM Arabic, launched by Stanford CRFM with Arabic.AI in December 2025, the Arabic-specialized models (AceGPT-v2, ALLaM, JAIS, SILMA) came out weaker than the multilingual ones, and the leaderboard’s authors attributed this to most of them being over a year old at evaluation time.

Put that beside the Falcon-H1-Arabic numbers and the real picture appears:

The wrong readingThe right reading
Arabic models are weaker than general onesOld Arabic models are weaker than new general ones
Specialization does not paySpecialization pays a lot, but model age outweighs it
Always pick the biggestPick the newest, then the most specialized, then the biggest

The practical lesson: model age is a stronger variable than people assume. A one year old specialized Arabic model loses to a three month old general model, and a brand new specialized Arabic model flips the equation back.


The practical verdict

  1. If you build inside the region and need local hosting, the specialized 7B is a serious option. It runs on modest hardware and its Arabic number clears models far larger.
  2. Do not read a leaderboard score as a dialect certificate. The leaderboard does not measure dialect at all, and the gap between a model that knows MSA and one that understands your user in Amman is the gap between a product that works and one that gets abandoned.
  3. Ask for the release date before you ask about size. Six months in this field is a generation.
  4. Make your system able to say “I am not sure.” Especially in dialect identification, where the social error costs more than the technical one.

Next in the series

Staying with Arabic models, I move to Jais 2 from Inception and MBZUAI, and the question most evaluators conflate: are Arab culture and Islamic knowledge the same thing? The PalmX 2025 numbers say clearly that they are not.


A new leaderboard trying to close the gap: QIMMA

Months after this article, the Technology Innovation Institute itself launched a new Arabic leaderboard called QIMMA, a serious attempt to address some of what I criticized above. I mention it here because criticism of leaderboards should be followed by acknowledgment when they improve.

ItemWhat QIMMA offers
Sample sizeMore than 52,000 samples
BenchmarksFourteen benchmarks across seven domains: culture, STEM, legal, medical, safety, poetry, coding
Native Arabic content99%
Sample validationMulti-model automated pipeline, then human review, discarding broken samples before evaluation
TransparencyPer-sample inference outputs published
NewThe first Arabic leaderboard to add code generation

Its first results in April 2026: Qwen3.5-397B-A17B at a mean of 68.06, then Karnak at 66.20, then Jais-2-70B-Chat at 65.81.

And the leaderboard’s own observation supports this article’s thesis directly: Arabic-specialized models lead on cultural and linguistic tasks, while multilingual models lead on code. Specialization pays where the language is the subject, not where the language is a container.

The leaderboard also exposed something about the benchmarks themselves: a 3.1% discard rate in ArabicMMLU due to answer quality and formatting problems. Three percent sounds small, and it approaches the size of the gaps between closely ranked models.

And what stayed outside the measurement

Even after QIMMA, the same cells remain empty:

measured today
│
unmeasured today
general knowledge
│
dialect generation
grammar
│
pragmatics and intent
content safety
│
multi-turn conversation
culture via MCQ
│
dialect stereotyping
poetry and code
│
speech and tone

The reason is the same every time: what a multiple-choice question can measure enters the leaderboards, and what requires a native speaker’s judgment stays outside. Which is why I say the gap is not in the models but in the measuring instruments, and closing it requires people rather than algorithms.


From the test notebook: five levels on one sentence

The leaderboard does not measure dialect, so this is what I measure instead. I start from a single sentence and build five levels on it:

“halla’ ma ili khili’ ahki maʿ hada.”

Level one, classification. Acceptable: Levantine family, medium to high confidence, no country named. Rejected: “Syrian” with confidence, because the same sentence is said in Beirut, Amman and Ramallah.

Level two, vocabulary in context. Khili’ here means mood, not creation, hada means anyone, and ma ili means I have no desire to. A model that renders khili’ as creation has failed at the first level.

Level three, one word in three contexts. Take ʿammar: a proper name in “ʿAmmar ija”, construction in “ʿammar al-beit khilis”, and a blessing in “bil-ʿammar ya rabb”. Three different answers or a fail.

Level four, structural conversion. Convert the sentence into Egyptian and Moroccan. Egyptian: “dilwa’ti mish fadi akallim hadd.” Moroccan: “daba ma ʿandish ma nahdar mʿa hatta wahed.” Notice the difference is not only lexical: the negation changed, the time adverb changed, and the whole sentence structure changed. A model returning three sentences with identical structure and swapped words did not convert a dialect, it swapped words.

Level five, register. The same meaning addressed to a friend, to a manager, and to an elder. The difference should show in length, directness and the apology device, not in the opening word alone.

The trap that specifically exposes Arabic models

I include this item for models trained on regional data, because it exposes the edges of that data:

“gal-li ya zalame shu sayir? gultillo wala ishi.”

The sentence mixes Bedouin or northern gal with Levantine zalame. A native speaker reads it immediately as a contact zone: northern Jordan, the Syrian badia, or a Jordanian of Palestinian origin in a mixed area.

A model trained on urban Levantine says “this is not Levantine”. One trained on Gulf data says “Gulf”. The correct answer is that it is Levantine with a Bedouin feature, and that dialects do not stop at political borders.

This single item measures something no Arabic leaderboard measures today: does the model know that dialect is a continuum rather than a closed box?

Sources

  • Technology Innovation Institute, Falcon-H1-Arabic announcement: falcon-lm.github.io/blog/falcon-h1-arabic
  • Open Arabic LLM Leaderboard v2 and its methodology: huggingface.co/blog/leaderboard-arabic-v2
  • HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
  • QIMMA Arabic leaderboard: huggingface.co/blog/tiiuae/qimma-arabic-leaderboard