The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Whisper large-v3, Modern Standard Arabic on the SADA set | 27.95% error | The model’s best case in Arabic | Open Universal Arabic ASR Leaderboard, Interspeech 2025 |
| Whisper large-v3, Najdi | 48.58% | The first stage of collapse | Same source |
| Whisper large-v3, Hijazi | 49.99% | Half the words wrong | Same source |
| Whisper large-v3, Egyptian | 59.28% | The most spoken Arabic dialect | Same source |
| Whisper large-v3, Khaliji | 59.92% | The worst, and roughly double | Same source |
| Arabic average across six test sets | 36.86% | Third of sixteen models | Same source |
| SeamlessM4T v2 for comparison | 38.16% | Fourth | Same source |
| What OpenAI published about Arabic | Nothing | Not a single figure | DevDay page and model card |
| What OpenAI published in general | A 10% to 20% error reduction against large-v2 | Across languages, with no breakdown | Model card on Hugging Face |
The gap: from 27.95 to 59.92, a factor of 2.1. And that is not a gap between two languages. It is a gap inside one language.
The error map by dialect
Read the bottom row the way your user reads it: six words in every ten come out wrong. No product is built on that.
Model card
| Item | Detail |
|---|---|
| Model | Whisper large-v3 |
| Lab | OpenAI |
| Release date | November 6, 2023, at DevDay |
| Licence | Open |
| What OpenAI announced | “Improved performance across languages” |
| Official Arabic numbers | None |
| Source of the numbers in this article | An academic paper from Elm in Saudi Arabia, Interspeech 2025 |
| The SADA set | Multidialectal Saudi Arabic speech data |
The ARAB-LENS reading: why this gap is more dangerous than any other
Across this series we have seen many gaps between Modern Standard Arabic and the dialects. But the speech gap is qualitatively different, for three reasons.
First: nobody speaks Modern Standard Arabic.
When you write, you might write formally. When you speak, you do not. Nobody says to a friend, in full Standard Arabic, “I would like to reserve an appointment at three o’clock”. They say it in Levantine or in Gulf. Which means the condition the model is good at, 27.95%, is the condition that almost never occurs in real use. And the condition that actually occurs is the one where it scores sixty percent error.
In text, formal Arabic is a common reality. In speech, formal Arabic is the exception.
Second: a speech error does not self-correct.
When a text model gets a word wrong, you read the sentence and understand what was meant. When a transcription system gets a word wrong, the word disappears and another one takes its place, and you do not know there was an error. The difference between “transferred one thousand” and “transferred two thousand” is not discovered until it is too late.
Third, and this is what bothers me most: the ranking is inverted.
Look at the order of the dialects in the table. The two worst are Egyptian and Khaliji. Egyptian is by a wide margin the most spoken Arabic dialect, and Gulf is the dialect of the region’s largest technology spending market. Which means the system is at its worst in the two dialects with the highest commercial value.
That is not random. It has a clear explanation: speech training data comes from the internet, and spoken Arabic online is mostly either news bulletins in formal Arabic or mixed entertainment content. Content that contains real everyday speech in one clean dialect with an accurate transcript is very rare.
And the final observation, which is a rule I keep repeating: OpenAI published not one Arabic figure for this model. The numbers above came from a Saudi team one year and nine months after release. I greatly respect that the most precise source on Arabic in this model is an Arab paper. But the question remains: why must we always measure for ourselves what these companies publish for every other language?
From the test notebook: building an Arabic speech test in an hour
This is a complete methodology you can run with your phone, and it gives you a real number for your product instead of a generic one.
Step one: collect twenty sentences from your own world.
Do not use generic sentences. Use the sentences your customer actually says. Examples from three sectors:
| Sector | A real sentence |
|---|---|
| Restaurant | “biddi talab delivery, talat shawarma djaaj w waahad lahme, w zidli toum.” |
| Pharmacy | “ʿandkom badeel la-had al-dawa? w bikam al-ʿilbe baʿd al-khasm?” |
| Bank | “hawwalt alfein w khamsmiyye min hsaabi imbaareh w ma wislat.” |
| Telecom | “al-baaqa khilsat w ana lissa bi-nuss al-shahr, shu al-hall?” |
| Bookings | “biddi aʾajjel al-hajz min al-khamees la-l-sabt, nafs al-ghurfe.” |
Step two: record them in three voices.
The same voice three times is not enough. You need a man’s voice, a woman’s voice, and the voice of someone over fifty. Transcription systems degrade noticeably with older voices, and that is a group nobody tests.
Step three: compute the error on what matters only.
Do not count every word. Count only the critical words:
| Category | Why it is critical |
|---|---|
| Numbers | Amounts, quantities, times |
| Names | People, products, places |
| Negation | “it did not arrive” against “it arrived” |
| Dates and days | Thursday against Saturday |
A system that misses 40% of words but gets every number and every negation right may be usable. A system that misses 20% but flips a negation is a catastrophe. Overall word error rate is a misleading metric for any real application.
Step four: test the four traps.
| The trap | Example | What it exposes |
|---|---|---|
| Short negation | “ma biddi” against “biddi” | One letter flips the meaning |
| Compound number | “alfein w khamsmiyye” | Number fragmentation |
| Uncommon name | “biddi ahki maʿ ustaz Muhannad” | Arabic names outside the common list |
| Foreign word | “el-appointment muʾajjal” | Code switching mid sentence |
Step five: repeat in six months.
Transcription systems improve fast, but not evenly across dialects. Keep your twenty recordings as a fixed set and rerun them. This becomes your own ruler, and it is more useful than any global leaderboard, because it measures your dialect and your customers.
Step six: measure the gender and age bias.
A dimension no Arabic leaderboard measures, and one I have watched sink entire projects:
| The speaker | What usually happens |
|---|---|
| Man, 25 to 45, common dialect | Best results |
| Woman, same age | Slight degradation |
| Over 60 | Heavy degradation |
| Child | Severe degradation |
| A regional accent within the dialect | Unpredictable |
If your product serves a government or health service, a high share of your audience is elderly. And measuring on your office colleagues’ voices does not represent them.
Why formal Arabic fails in speech to begin with
I want to explain something that gets misread constantly: even the “good” figure, 27.95%, is not good.
| The comparison | Approximate error rate |
|---|---|
| English in clean conditions | Under 5% |
| Modern Standard Arabic, this model | 27.95% |
| Gulf Arabic, this model | 59.92% |
Look at the first and second rows. The best case in Arabic is roughly five times worse than the ordinary case in English. And that is before dialects enter the picture at all.
I raise this because the Arabic conversation dwells heavily on the formal-versus-dialect gap and forgets the larger one: the gap between Arabic as a whole and English. The first is an internal gap we can close with local data. The second is a structural gap that needs investment of a different order.
The practical verdict
If you are building an Arabic voice assistant, do not make voice the only path. Route every financial or critical decision through a text confirmation. That is not weak design. It is engineering that respects a sixty percent error rate.
And do not believe a single number for “Arabic”. Ask any vendor for a per-dialect breakdown, and anyone who does not have one has not tested your product.
If your audience is Gulf or Egyptian, know that you are in the two hardest cases specifically, and budget for correction and human review from the start.
Next in the series
One month later the French lab Mistral released Mixtral 8x7B and announced the languages it masters: English, French, Italian, German and Spanish. Five languages. The next article is about the list Arabic was not on, and what it means to be excluded explicitly rather than by omission.
Sources
- OpenAI, New models and developer products announced at DevDay, November 6, 2023: openai.com
- Whisper release v20231106 on GitHub: github.com/openai/whisper/releases
- Wang, Alhmoud and Alqurishi, Open Universal Arabic ASR Leaderboard, Interspeech 2025, August 2025: isca-archive.org
- Model card openai/whisper-large-v3 on Hugging Face