The numbers first
| System | Word error rate | Leaderboard or source |
|---|---|---|
| Cohere Transcribe Arabic | 25.87 | Hugging Face Arabic ASR Leaderboard, July 2026 |
| Meta OmniASR-LLM-7B | 2.45 points behind the leader | Same leaderboard |
| OpenAI Whisper Large V3 | About 11 points behind the leader | Same leaderboard |
| Munsit-1 | 26.68 | Open Universal Arabic ASR Leaderboard |
| OpenAI Whisper on that same board | 36.86 | Open Universal Arabic ASR Leaderboard |
| NADI 2025, best participating system | 35.68 WER, 12.20 CER | ArabicNLP 2025 |
| NADI 2025, spoken dialect identification | 79.8% | ArabicNLP 2025 |
| NADI 2025, diacritic restoration | 55 WER, 13 CER | ArabicNLP 2025 |
| AHELM, best audio-language model | 0.803 mean win rate | Stanford CRFM |
| AHELM, baseline (ASR then LLM) | 6th of 14, and more robust to noise | Stanford CRFM |
A quarter of the words. That is the state of the art in multidialectal Arabic transcription today. If you are building an Arabic voice product, that number, not a marketing page, should govern your architectural decisions.
A methodological warning before any comparison
Before anyone drops that table into a slide deck, here is what the researchers themselves say: WER figures are not comparable across vendors except with great caution, because test sets differ, normalization rules differ, and the handling of diacritics, hamza and punctuation differs.
Look at Whisper in the table: it appears once as eleven points behind a leader on one board, and once at 36.86 on another. Same model. The difference is in the measurement, not the system.
So ARAB-LENS carries a rule: compare only within a single leaderboard, and treat cross-board figures as directional signals rather than verdicts.
Model card for the leader
| Field | Detail |
|---|---|
| Name | Cohere Transcribe Arabic |
| Lab | Cohere |
| Released | July 7, 2026 |
| Base | A two billion parameter ASR model |
| Coverage | About 30 Arabic varieties: MSA, Egyptian, Gulf, Levantine, Maghrebi |
| Extra capability | Arabic to English code-switching, and English spoken with an Arabic accent |
| Human preference | Preferred over Whisper in 95.8% of human tests |
| Speed | RTFx 525 against 146 for Whisper |
| License | Apache 2.0 |
That 95.8% human preference matters more to me than the 25.87, because it measures what a user feels rather than what a script computes. Which is exactly what I argue for in every article: the final verdict belongs to a human ear.
Why WER alone is an incomplete metric in Arabic
Four reasons, all structural in the language rather than the model:
One, errors do not weigh the same. Ma ba’dar against ba’dar: one word in the arithmetic, a total inversion of meaning. Biddi ruh against biddi aruh: an error in the arithmetic, zero error in meaning.
Two, normalization moves the number. Does hamza count? Ta marbuta? Diacritics? Changing one rule shifts WER by several points without the system changing at all.
Three, diacritics are a task of their own. NADI 2025 recorded 55 WER on diacritic restoration for speech, meaning more than half the words carried wrong vowelling. Anyone building an educational or Quranic product cannot ignore that figure.
Four, transcription is not comprehension. The most important one. A system that transcribes accurately can still fail a simple question about what was said.
The test: a four-stage protocol for your team
Stage A, the reality number. Record ten minutes of your actual users, not a trained reader. Five samples in a quiet room, five in the real world: a cafe, a car, a phone call. Compute WER for both sets. The gap between them is what you will pay in production.
Stage B, names and numbers. Put a compound Arabic proper name, a long spoken number, and a code switch into every recording. These are the three places systems break, and they are exactly what your product needs: customer names, amounts, technical terms.
Stage C, comprehension over copying. After transcription, ask the language model a question about the meaning rather than the text. “Why did the speaker not come?” and “was he satisfied?”
Stage D, dialect and confidence. Ask for a dialect label with a confidence level, and accept “Levantine” instead of “Syrian.” Then watch: does the system admit uncertainty, or does it always guess a country?
The architectural decision: one path or two
This is the most practical section here.
| Integrated audio model | Transcribe then reason | |
|---|---|---|
| Speed | Faster, one step | Slower, two steps |
| Demo quality | Impressive | Ordinary |
| Human review | Hard, no intermediate text | Easy, the text exists |
| Correction | Requires rerunning the whole turn | Fix the text, rerun the understanding |
| Noise robustness | Lower, per AHELM | Higher, per AHELM |
| Audit and compliance | Harder | Easier, there is a written record |
The AHELM evaluation supports this directly: the simple baseline, transcription followed by a language model, placed sixth among fourteen systems and was more robust to noise than most of them.
My recommendation for anyone building in Arabic today: start with two paths, not one. Not because the integrated model is weak, but because the intermediate text is worth more in Arabic than it appears. You need it for review, for dialect correction, and for training your own model on your own data later.
The practical verdict, and the series conclusion
- Treat 25% error as the starting point, not the destination. Design your interface assuming transcription will be wrong: show the text, make correcting it one tap.
- Measure in your environment and your dialect. A public leaderboard tells you nothing about a cafe in Amman.
- Separate transcription from understanding. The reason is not only technical, it is operational and regulatory.
- Accept uncertainty in dialect, and put correction in the user’s hands.
That closes the ARAB-LENS series at eighteen articles. The thread running through all of them is one sentence: Arabic is measured today where measurement is easy, not where the user lives. MSA is measured, dialect is not. Knowledge is measured, intent is not. Transcription is measured, comprehension is not. And everyone building an Arabic product pays for that gap, whether they notice it or not.
My next step is a Levantine gold set written by native speakers, withheld from training, used as a reference rather than as training material. Anyone who wants to help build it knows where to find me.
What no number in this article measures
Every figure above concerns transcription, turning sound into characters. That is half of a voice product. The other half is the system understanding what was said, and answering in a voice that sounds human.
On understanding specifically there is a published number worth reading alongside the WER figures: the AL-QASIDA study, which measured nine text models across eight dialects, found that understanding outpaces generation, and that models fall back to MSA with a 0.99 correlation when the dialect is beyond them.
In a voice product that means this: even if your system transcribes Levantine speech excellently, the layer that answers the user will drift toward MSA. You end up with an assistant that hears you in your dialect and speaks to you in broadcast language.
The pre-launch checklist
| Item | Acceptance criterion |
|---|---|
| WER in a quiet environment | Measure it, do not accept the vendor’s number |
| WER in a real environment | Less than ten points worse than quiet |
| Proper names | 90% correct across a sample of fifty names |
| Spoken numbers | Correct in form, not read digit by digit |
| Code-switching | Does not translate technical terms |
| Understanding intent | Answers a question about meaning, not about text |
| Dialect identification | A family with a confidence level, not a definitive country |
| Speech generation | Correct pausal form on ta marbuta, correct interrogative intonation |
| Correction | The user can edit the text in one tap |
| Logging | Correction rate stored by dialect and environment |
The last item is what turns your product into a system that learns: user correction rates broken down by dialect and environment give you, within three months, the real weakness map of your system, and no company sells that map and no leaderboard publishes it.
Closing the series: what I am asking of the field
Eighteen articles, one thread: we measure Arabic where measurement is easy, not where the user lives.
I close with three specific asks, all of which I consider achievable within a year:
- A public dialect generation benchmark judged by native speakers. The tooling exists. What is needed is funding for human hours, not a new algorithm.
- A Levantine gold set withheld from training. Because Levantine is the weakest in every available measurement and the most absent from dialogue benchmarks.
- Publishing a fluent-and-wrong rate with every Arabic evaluation. One number saying how often an answer looks right and is not.
Anyone who wants to work on any of the three knows where to find me.
From the speech notebook: the fixed sample I run
My audio sample has been fixed for years, and this is exactly what is in it:
| Part | Content | What it exposes |
|---|---|---|
| 1 | Calm MSA reading, ten seconds | The baseline, any system passes here |
| 2 | Ordinary Damascene Levantine, twenty seconds | Common urban representation |
| 3 | Fast Levantine with assimilation | Real speech rather than read speech |
| 4 | Northern Jordanian with g for qaf | Outside urban representation |
| 5 | A sentence with a compound name and a long number | The most common breaking points |
| 6 | The same sentence in a cafe | The gap between lab and reality |
And the two numbers I log per part: word error rate, and meaning-changing error rate. The second is what I base decisions on.
Three transcription errors I see repeatedly, and why they matter
One, the negative particle disappears. Ma ba’dar becomes ba’dar. One word in the arithmetic, a total inversion of meaning. In a support product that means the system recorded an agreement instead of a refusal.
Two, compound names break apart. ʿAbd al-Rahman becomes two unrelated words. A small error in the arithmetic, and an unmatchable name in a customer database.
Three, numbers invert. Alfein w khamsmiyye comes out as 2500 or 5002 depending on the system, because Arabic speaks units before tens in some forms. In a financial application that error is intolerable.
These three are what I check by hand in every evaluation, because any averaged metric swallows them.
Speech generation test: five sentences and a human verdict
| Sentence | What it measures |
|---|---|
| “safart min ʿAmman ila ʿUman” | Distinguishing two names written almost identically |
| “qara’tu risalata al-mudira” | Pausal ta marbuta as an h |
| “al-mablagh alfan wa thalathumi’a wa khamsa wa sabʿun” | Speaking the number in form, not digit by digit |
| “inta rayih la-halak?” | Interrogative intonation |
| “hajazt ʿala tayaran Qatar Airways” | Transition between two sound systems |
And the final verdict goes to five Arab listeners, with one question: does this sound like a person or like a system?
The series conclusion
Eighteen articles, one thread: we measure Arabic where measurement is easy, not where the user lives.
What I ask of the field is three specific things: a public dialect generation benchmark judged by native speakers, a Levantine gold set withheld from training, and a published fluent-and-wrong rate with every Arabic evaluation. All three are achievable within a year, none of them needs a new algorithm, only human hours and a decision to measure what matters rather than what is easy.
Sources
- Cohere, Transcribe Arabic launch and results: cohere.com/blog/transcribe-arabic
- Arabic speech-to-text tool comparison, Munsit and Whisper figures: munsit.com/blog/best-arabic-speech-to-text
- NADI 2025, first multidialectal Arabic speech processing shared task: arxiv.org/abs/2509.02038
- AHELM, holistic evaluation of audio-language models, Stanford CRFM: arxiv.org/pdf/2508.21376
- NADI 2026, the upcoming mixed-dialect Arabic speech task: nadi.dlnlp.ai/2026
- Practitioner comparison of Arabic transcription tools: munsit.com/blog/best-arabic-speech-to-text