A Trillion Arabic Phonetic Segments: Reading Fanar 2.0, the First Arabic Multimodal Stack

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: December 9, 2025 | Published: December 26, 2025

The numbers first

ItemNumberSource
Model size27B parameters, triple the original Fanar’s 9BQatar Computing Research Institute, December 2025
Text training dataMore than 300 billion wordsSame source
Speech training dataMore than one trillion Arabic phonetic segmentsSame source
ModalitiesText, image, speech to text, text to speech, audio understanding, Arabic poetry generationSame source
Safety layerFanarGuard, a moderation filter detecting safety violations and cultural misalignment in Arabic and EnglishSame source
Next releaseFanar 3.0, planned for December 2026Statement by Dr. Ahmed Elmagarmid
Fanar-1-9B in PalmX 2025Among the most used models by participating teamsArabicNLP 2025
Competitor comparisonFalcon-H1-Arabic 7B claims to surpass Fanar-9B on OALLTechnology Innovation Institute, January 2026

The figure that stopped me in this list is not the 27 billion parameters or the 300 billion words. It is one trillion phonetic segments. Because the biggest obstacle in front of Arabic AI is not text, it is speech. Arabic text is abundant online. Arabic speech, accurately transcribed and labeled across dialects, is rare and expensive and requires humans who know the difference between a pronunciation error and a regional variant. Whoever assembles a trillion Arabic phonetic segments is building infrastructure, not a model.


Why that infrastructure matters

Put the Fanar figures beside the hardest published Arabic speech task, NADI 2025:

TaskBest NADI 2025 resultReading
Speech recognition35.68 WER, 12.20 CERMore than a third of spoken words come out wrong
Spoken dialect identification79.8% accuracyAcceptable for a family, not for a region
Diacritic restoration for speech55 WER, 13 CERMore than half the words are wrongly vowelled
Teams actually submitting8 of 44 registeredA small research field relative to the size of the language
The state of Arabic speech in 2025
Dialect transcription
64% of words correct
Dialect identification
79.8%
Spoken diacritics
about 45%
Understanding intent
not measured at all

In that context, announcing an Arabic model that combines transcription, speech generation and understanding in one stack is not a marketing item, it is an attempt to close a real gap. And the question still standing: there are no published third-party numbers for Fanar 2.0’s speech performance. What we have is a description of capabilities and a data volume, not an error rate.


Model card

FieldDetail
NameFanar 2.0
LabQatar Computing Research Institute, Hamad Bin Khalifa University
AnnouncedWorld Summit AI in Doha, December 9, 2025
Size27B parameters
Distinguishing featureFirst Arabic stack combining text, image and speech with a cultural safety layer
Axis under testARAB-SPEECH and ARAB-GEN together

Testing the speech stack in four stages

Stage A, transcription. Record thirty seconds in your dialect containing a compound proper name, a spoken number, and a code switch. Compute WER yourself, then record the same thing in a cafe and measure the difference. The gap between the two is your “reality number,” and it is the only one that matters.

Stage B, understanding. Play the sentence “wallah kint biddi aji mbarih, bas sar maʿi shaghle w ma ‘dirt,” and ask for the reason without requesting a transcript.

Stage C, speech generation. This is the side everyone neglects. Ask the model to speak:

  • A sentence with a foreign name inside Arabic: “hajazt tadhkara ʿala tayaran Qatar Airways al-saʿa sabʿa.”
  • A sentence with a long number: “al-mablagh alfein w thalathmiyye w khamsa w sabʿin riyalan.”
  • A sentence with a final hamza and a ta marbuta in both connected and pausal form: “qara’tu risalata al-mudira.”
  • A question with genuine interrogative intonation rather than flat delivery: “inta rayih la-halak?”

The judging criterion is not “is the audio clear.” It is does this sound natural to an Arab ear: is the intonation right on the question, does the pause on the ta marbuta turn it into a ha, is the number spoken in its correct form rather than read digit by digit.

Stage D, cultural safety. The FanarGuard layer claims to detect “cultural misalignment,” and that claim deserves a serious test rather than applause:

  • Give it a text mixing a Moroccan custom with a Gulf one and see whether it catches the blend.
  • Give it a sentence stereotyping the people of one region and see whether it flags it.
  • Then the important one: give it a perfectly acceptable text that uses a common Levantine colloquial expression, and see whether it wrongly flags that. Over-moderation of culture is a defect, not a feature, because it will erase your user’s dialect in the name of protecting them.

A methodological note on cross comparisons

In January 2026 the Technology Innovation Institute announced that Falcon-H1-Arabic 7B surpasses Fanar-9B on the OALL leaderboard. The number is accurate as published, but reading it requires three caveats:

  1. The comparison is against Fanar-1-9B, not Fanar 2.0, which is 27B and shipped a month before that announcement.
  2. The figures come from a competitor, which does not make them wrong, it means the choice of what to compare against is also a marketing decision.
  3. OALL does not measure speech at all, and speech is what Fanar 2.0 was built around. Comparing a text model and a multimodal stack on a text leaderboard measures half of one of them.

This is not a defense of one model against another. It is a general rule for reading numbers in this field: always ask who published the number, on which version, on which leaderboard, and does that leaderboard measure the thing the model claims to lead in?


The practical verdict

  1. If your product is Arabic and voice-based, specialized Arabic speech models deserve a serious trial, not because they are necessarily more accurate, but because their speech training data is natively Arabic rather than translated.
  2. Measure on your users, not on a clean sample. The distance between a studio and a cafe is larger in Arabic than in English, because noise swallows the short vowels.
  3. Test the safety layer in both directions. What it catches, and what it catches wrongly. A filter that treats dialect as deviation costs you a user who does not know why their words were rejected.
  4. Do not wait for third-party numbers to decide. As of this writing there are no independent figures for Fanar 2.0’s speech performance, and ten minutes of recording from your own environment tells you what no leaderboard will.

Next in the series

To Saudi Arabia and ALLaM, carrying the most cited number in Arabic evaluation: the top of ArabicMMLU. And an uncomfortable question: is knowing school curricula in Arabic evidence of understanding Arabic?


What the market actually offers an Arabic voice builder

Setting research papers aside, here is the commercially available picture today, as published practitioner comparisons describe it. I include it because a team building now is not waiting for a paper, it is picking a vendor this week.

SystemDialect coverage as advertised
MunsitMore than 25 dialects, no manual dialect setting
SpeechmaticsGulf, Egyptian, Levantine, Maghrebi, with code-switching
Deepgram Nova-317 Arabic variants across Gulf, MSA, Egyptian, Levantine
WhisperMSA-leaning, limited dialectal generalization
Amazon TranscribeGulf and MSA only

And the WER figures published in that same comparison: Munsit-1 at 26.68% against Whisper at 36.86% on multidialectal test sets.

The most honest part of that comparison, though, is its warning: WER figures are not comparable across vendors, because test sets and normalization rules differ. Which means the table above is useful for shortening your shortlist, not for making the decision.

The build or buy question

SituationThe sensible decision
One dialect, small volumeA ready vendor, measured on your own samples
Many dialects, large volumeA vendor first, then fine-tune an open model on your data once you have enough hours
Sensitive data that cannot leave the countryA locally hosted open model, even at lower accuracy initially
Religious or educational content needing diacriticsNever hand automatic vowelling to the user, add human review
High-noise environmentsThe separate transcription path, for its observed noise robustness

The item everyone skips: speech generation

Almost all published discussion is about transcription, while half of a voice product is the spoken reply. I measure four things in Arabic speech generation, and no published number measures any of them:

  1. Pausing on ta marbuta as a ha, not a ta.
  2. Pronouncing compound numbers in their correct form rather than digit by digit.
  3. Interrogative intonation, since a spoken Arabic question changes at its end.
  4. Foreign names inside an Arabic sentence, the single clearest giveaway of a synthetic voice.

Five sentences containing all four are enough to judge, and five Arab listeners settle it in five minutes. That simple human measurement is more accurate than any automatic metric available today for Arabic speech generation.


From the speech notebook: five sentences that expose any Arabic voice generator

I use these five to evaluate speech generation rather than transcription, the side no leaderboard measures:

Sentence one, distinguishing ʿAmman from ʿUman.

“safart min ʿAmman ila ʿUman.”

The first is Jordan’s capital, with a fatha and a doubled m. The second is the Sultanate, with a damma and no doubling. A generator that pronounces them identically reveals that it is reading letters rather than knowing words. This is my fastest test of all.

Sentence two, pausal ta marbuta.

“qara’tu risalata al-mudira.”

The pause falls on an h: al-mudirah. Pronouncing an open t at the end of the sentence is an error no native speaker makes.

Sentence three, the compound number.

“al-mablagh alfan wa thalathumi’a wa khamsa wa sabʿun riyalan.”

Arabic speaks the units before the tens, and a weak generator reads the figure digit by digit or inverts the order.

Sentence four, interrogative intonation.

“inta rayih la-halak?”

A question in spoken Arabic often carries no interrogative particle, and its only marker is a rise at the end. A generator that delivers it flat converts the question into a statement, and that single failure is what makes synthetic voices sound synthetic.

Sentence five, a foreign name inside Arabic.

“hajazt ʿala tayaran Qatar Airways al-saʿa sabʿa.”

What is required is a natural transition between the two sound systems, neither an exaggerated English accent nor a forced Arabization.

How I judge, and who judges

I do not judge these sentences with my ear alone. I play them for five Arab listeners from different dialects and ask one question only: does this sound like a person or like a system? Then I compute the share who said system.

The reason for that simple procedure is that automatic audio quality metrics measure clarity and cleanliness, not naturalness. I have heard Arabic voices that were perfectly clean and completely grating, because the intonation was English and the words were Arabic.

A note on cultural filtering

A cultural safety layer needs testing in both directions. Send a perfectly acceptable text containing a common Levantine expression such as “ya zalame shu sayir” or “Allah yirda ʿaleik”, and see whether the filter flags it as a violation. Over-moderating dialect is not protection, it is deleting the user’s voice, and it is a defect that should be logged like any error.

Sources

  • Fanar 2.0 announcement at World Summit AI, Doha: middleeastainews.com
  • PalmX 2025 and the use of Fanar-1-9B by participating teams: arxiv.org/abs/2509.02550
  • NADI 2025 Arabic multidialectal speech numbers: arxiv.org/abs/2509.02038
  • Technology Innovation Institute, Falcon-H1-Arabic comparisons: falcon-lm.github.io/blog/falcon-h1-arabic
  • Arabic practitioner comparisons of speech models: munsit.com and kuwaitai.net