25.87: The Best Number in Arabic Speech Transcription, and What It Actually Means

Reading Time: 10 min
18
Model released: July 7, 2026 | Published: July 21, 2026 Ehab Saleh | techkahwa.net

The numbers first

SystemWord error rateLeaderboard or source
Cohere Transcribe Arabic25.87Hugging Face Arabic ASR Leaderboard, July 2026
Meta OmniASR-LLM-7B2.45 points behind the leaderSame leaderboard
OpenAI Whisper Large V3About 11 points behind the leaderSame leaderboard
Munsit-126.68Open Universal Arabic ASR Leaderboard
OpenAI Whisper on that same board36.86Open Universal Arabic ASR Leaderboard
NADI 2025, best participating system35.68 WER, 12.20 CERArabicNLP 2025
NADI 2025, spoken dialect identification79.8%ArabicNLP 2025
NADI 2025, diacritic restoration55 WER, 13 CERArabicNLP 2025
AHELM, best audio-language model0.803 mean win rateStanford CRFM
AHELM, baseline (ASR then LLM)6th of 14, and more robust to noiseStanford CRFM
Arabic transcription: where we actually are
Cohere Transcribe Arabic
25.87 WER
Munsit-1
26.68 WER
NADI 2025 best entry
35.68 WER
Whisper Large V3
36.86 WER
Lower is better. And even the best misses a quarter of the words

A quarter of the words. That is the state of the art in multidialectal Arabic transcription today. If you are building an Arabic voice product, that number, not a marketing page, should govern your architectural decisions.


A methodological warning before any comparison

Before anyone drops that table into a slide deck, here is what the researchers themselves say: WER figures are not comparable across vendors except with great caution, because test sets differ, normalization rules differ, and the handling of diacritics, hamza and punctuation differs.

Look at Whisper in the table: it appears once as eleven points behind a leader on one board, and once at 36.86 on another. Same model. The difference is in the measurement, not the system.

So ARAB-LENS carries a rule: compare only within a single leaderboard, and treat cross-board figures as directional signals rather than verdicts.


Model card for the leader

FieldDetail
NameCohere Transcribe Arabic
LabCohere
ReleasedJuly 7, 2026
BaseA two billion parameter ASR model
CoverageAbout 30 Arabic varieties: MSA, Egyptian, Gulf, Levantine, Maghrebi
Extra capabilityArabic to English code-switching, and English spoken with an Arabic accent
Human preferencePreferred over Whisper in 95.8% of human tests
SpeedRTFx 525 against 146 for Whisper
LicenseApache 2.0

That 95.8% human preference matters more to me than the 25.87, because it measures what a user feels rather than what a script computes. Which is exactly what I argue for in every article: the final verdict belongs to a human ear.


Why WER alone is an incomplete metric in Arabic

Four reasons, all structural in the language rather than the model:

One, errors do not weigh the same. Ma ba’dar against ba’dar: one word in the arithmetic, a total inversion of meaning. Biddi ruh against biddi aruh: an error in the arithmetic, zero error in meaning.

Two, normalization moves the number. Does hamza count? Ta marbuta? Diacritics? Changing one rule shifts WER by several points without the system changing at all.

Three, diacritics are a task of their own. NADI 2025 recorded 55 WER on diacritic restoration for speech, meaning more than half the words carried wrong vowelling. Anyone building an educational or Quranic product cannot ignore that figure.

Four, transcription is not comprehension. The most important one. A system that transcribes accurately can still fail a simple question about what was said.


The test: a four-stage protocol for your team

Stage A, the reality number. Record ten minutes of your actual users, not a trained reader. Five samples in a quiet room, five in the real world: a cafe, a car, a phone call. Compute WER for both sets. The gap between them is what you will pay in production.

Stage B, names and numbers. Put a compound Arabic proper name, a long spoken number, and a code switch into every recording. These are the three places systems break, and they are exactly what your product needs: customer names, amounts, technical terms.

Stage C, comprehension over copying. After transcription, ask the language model a question about the meaning rather than the text. “Why did the speaker not come?” and “was he satisfied?”

Stage D, dialect and confidence. Ask for a dialect label with a confidence level, and accept “Levantine” instead of “Syrian.” Then watch: does the system admit uncertainty, or does it always guess a country?


The architectural decision: one path or two

This is the most practical section here.

Integrated audio modelTranscribe then reason
SpeedFaster, one stepSlower, two steps
Demo qualityImpressiveOrdinary
Human reviewHard, no intermediate textEasy, the text exists
CorrectionRequires rerunning the whole turnFix the text, rerun the understanding
Noise robustnessLower, per AHELMHigher, per AHELM
Audit and complianceHarderEasier, there is a written record

The AHELM evaluation supports this directly: the simple baseline, transcription followed by a language model, placed sixth among fourteen systems and was more robust to noise than most of them.

My recommendation for anyone building in Arabic today: start with two paths, not one. Not because the integrated model is weak, but because the intermediate text is worth more in Arabic than it appears. You need it for review, for dialect correction, and for training your own model on your own data later.


The practical verdict, and the series conclusion

  1. Treat 25% error as the starting point, not the destination. Design your interface assuming transcription will be wrong: show the text, make correcting it one tap.
  2. Measure in your environment and your dialect. A public leaderboard tells you nothing about a cafe in Amman.
  3. Separate transcription from understanding. The reason is not only technical, it is operational and regulatory.
  4. Accept uncertainty in dialect, and put correction in the user’s hands.

That closes the ARAB-LENS series at eighteen articles. The thread running through all of them is one sentence: Arabic is measured today where measurement is easy, not where the user lives. MSA is measured, dialect is not. Knowledge is measured, intent is not. Transcription is measured, comprehension is not. And everyone building an Arabic product pays for that gap, whether they notice it or not.

My next step is a Levantine gold set written by native speakers, withheld from training, used as a reference rather than as training material. Anyone who wants to help build it knows where to find me.


What no number in this article measures

Every figure above concerns transcription, turning sound into characters. That is half of a voice product. The other half is the system understanding what was said, and answering in a voice that sounds human.

On understanding specifically there is a published number worth reading alongside the WER figures: the AL-QASIDA study, which measured nine text models across eight dialects, found that understanding outpaces generation, and that models fall back to MSA with a 0.99 correlation when the dialect is beyond them.

In a voice product that means this: even if your system transcribes Levantine speech excellently, the layer that answers the user will drift toward MSA. You end up with an assistant that hears you in your dialect and speaks to you in broadcast language.

The pre-launch checklist

ItemAcceptance criterion
WER in a quiet environmentMeasure it, do not accept the vendor’s number
WER in a real environmentLess than ten points worse than quiet
Proper names90% correct across a sample of fifty names
Spoken numbersCorrect in form, not read digit by digit
Code-switchingDoes not translate technical terms
Understanding intentAnswers a question about meaning, not about text
Dialect identificationA family with a confidence level, not a definitive country
Speech generationCorrect pausal form on ta marbuta, correct interrogative intonation
CorrectionThe user can edit the text in one tap
LoggingCorrection rate stored by dialect and environment

The last item is what turns your product into a system that learns: user correction rates broken down by dialect and environment give you, within three months, the real weakness map of your system, and no company sells that map and no leaderboard publishes it.

Closing the series: what I am asking of the field

Eighteen articles, one thread: we measure Arabic where measurement is easy, not where the user lives.

I close with three specific asks, all of which I consider achievable within a year:

  1. A public dialect generation benchmark judged by native speakers. The tooling exists. What is needed is funding for human hours, not a new algorithm.
  2. A Levantine gold set withheld from training. Because Levantine is the weakest in every available measurement and the most absent from dialogue benchmarks.
  3. Publishing a fluent-and-wrong rate with every Arabic evaluation. One number saying how often an answer looks right and is not.

Anyone who wants to work on any of the three knows where to find me.


From the speech notebook: the fixed sample I run

My audio sample has been fixed for years, and this is exactly what is in it:

PartContentWhat it exposes
1Calm MSA reading, ten secondsThe baseline, any system passes here
2Ordinary Damascene Levantine, twenty secondsCommon urban representation
3Fast Levantine with assimilationReal speech rather than read speech
4Northern Jordanian with g for qafOutside urban representation
5A sentence with a compound name and a long numberThe most common breaking points
6The same sentence in a cafeThe gap between lab and reality

And the two numbers I log per part: word error rate, and meaning-changing error rate. The second is what I base decisions on.

Three transcription errors I see repeatedly, and why they matter

One, the negative particle disappears. Ma ba’dar becomes ba’dar. One word in the arithmetic, a total inversion of meaning. In a support product that means the system recorded an agreement instead of a refusal.

Two, compound names break apart. ʿAbd al-Rahman becomes two unrelated words. A small error in the arithmetic, and an unmatchable name in a customer database.

Three, numbers invert. Alfein w khamsmiyye comes out as 2500 or 5002 depending on the system, because Arabic speaks units before tens in some forms. In a financial application that error is intolerable.

These three are what I check by hand in every evaluation, because any averaged metric swallows them.

Speech generation test: five sentences and a human verdict

SentenceWhat it measures
“safart min ʿAmman ila ʿUman”Distinguishing two names written almost identically
“qara’tu risalata al-mudira”Pausal ta marbuta as an h
“al-mablagh alfan wa thalathumi’a wa khamsa wa sabʿun”Speaking the number in form, not digit by digit
“inta rayih la-halak?”Interrogative intonation
“hajazt ʿala tayaran Qatar Airways”Transition between two sound systems

And the final verdict goes to five Arab listeners, with one question: does this sound like a person or like a system?

The series conclusion

Eighteen articles, one thread: we measure Arabic where measurement is easy, not where the user lives.

What I ask of the field is three specific things: a public dialect generation benchmark judged by native speakers, a Levantine gold set withheld from training, and a published fluent-and-wrong rate with every Arabic evaluation. All three are achievable within a year, none of them needs a new algorithm, only human hours and a decision to measure what matters rather than what is easy.

Sources

  • Cohere, Transcribe Arabic launch and results: cohere.com/blog/transcribe-arabic
  • Arabic speech-to-text tool comparison, Munsit and Whisper figures: munsit.com/blog/best-arabic-speech-to-text
  • NADI 2025, first multidialectal Arabic speech processing shared task: arxiv.org/abs/2509.02038
  • AHELM, holistic evaluation of audio-language models, Stanford CRFM: arxiv.org/pdf/2508.21376
  • NADI 2026, the upcoming mixed-dialect Arabic speech task: nadi.dlnlp.ai/2026
  • Practitioner comparison of Arabic transcription tools: munsit.com/blog/best-arabic-speech-to-text