27.95 Against 59.92: The Factor of Two That Separates Your Formal Arabic From Your Dialect

Reading Time: 8 min
18
Ehab Saleh | techkahwa.net Model released: November 6, 2023 | Published: November 20, 2023

The numbers first

Metric Result What it measures Source
Whisper large-v3, Modern Standard Arabic on the SADA set 27.95% error The model’s best case in Arabic Open Universal Arabic ASR Leaderboard, Interspeech 2025
Whisper large-v3, Najdi 48.58% The first stage of collapse Same source
Whisper large-v3, Hijazi 49.99% Half the words wrong Same source
Whisper large-v3, Egyptian 59.28% The most spoken Arabic dialect Same source
Whisper large-v3, Khaliji 59.92% The worst, and roughly double Same source
Arabic average across six test sets 36.86% Third of sixteen models Same source
SeamlessM4T v2 for comparison 38.16% Fourth Same source
What OpenAI published about Arabic Nothing Not a single figure DevDay page and model card
What OpenAI published in general A 10% to 20% error reduction against large-v2 Across languages, with no breakdown Model card on Hugging Face

The gap: from 27.95 to 59.92, a factor of 2.1. And that is not a gap between two languages. It is a gap inside one language.


The error map by dialect

Word error rate for Whisper large-v3 on the SADA set
MSA
27.95%
Najdi
48.58%
Hijazi
49.99%
Egyptian
59.28%
Khaliji
59.92%

Read the bottom row the way your user reads it: six words in every ten come out wrong. No product is built on that.


Model card

Item Detail
Model Whisper large-v3
Lab OpenAI
Release date November 6, 2023, at DevDay
Licence Open
What OpenAI announced “Improved performance across languages”
Official Arabic numbers None
Source of the numbers in this article An academic paper from Elm in Saudi Arabia, Interspeech 2025
The SADA set Multidialectal Saudi Arabic speech data

The ARAB-LENS reading: why this gap is more dangerous than any other

Across this series we have seen many gaps between Modern Standard Arabic and the dialects. But the speech gap is qualitatively different, for three reasons.

First: nobody speaks Modern Standard Arabic.

When you write, you might write formally. When you speak, you do not. Nobody says to a friend, in full Standard Arabic, “I would like to reserve an appointment at three o’clock”. They say it in Levantine or in Gulf. Which means the condition the model is good at, 27.95%, is the condition that almost never occurs in real use. And the condition that actually occurs is the one where it scores sixty percent error.

In text, formal Arabic is a common reality. In speech, formal Arabic is the exception.

Second: a speech error does not self-correct.

When a text model gets a word wrong, you read the sentence and understand what was meant. When a transcription system gets a word wrong, the word disappears and another one takes its place, and you do not know there was an error. The difference between “transferred one thousand” and “transferred two thousand” is not discovered until it is too late.

Third, and this is what bothers me most: the ranking is inverted.

Look at the order of the dialects in the table. The two worst are Egyptian and Khaliji. Egyptian is by a wide margin the most spoken Arabic dialect, and Gulf is the dialect of the region’s largest technology spending market. Which means the system is at its worst in the two dialects with the highest commercial value.

That is not random. It has a clear explanation: speech training data comes from the internet, and spoken Arabic online is mostly either news bulletins in formal Arabic or mixed entertainment content. Content that contains real everyday speech in one clean dialect with an accurate transcript is very rare.

And the final observation, which is a rule I keep repeating: OpenAI published not one Arabic figure for this model. The numbers above came from a Saudi team one year and nine months after release. I greatly respect that the most precise source on Arabic in this model is an Arab paper. But the question remains: why must we always measure for ourselves what these companies publish for every other language?


From the test notebook: building an Arabic speech test in an hour

This is a complete methodology you can run with your phone, and it gives you a real number for your product instead of a generic one.

Step one: collect twenty sentences from your own world.

Do not use generic sentences. Use the sentences your customer actually says. Examples from three sectors:

Sector A real sentence
Restaurant “biddi talab delivery, talat shawarma djaaj w waahad lahme, w zidli toum.”
Pharmacy “ʿandkom badeel la-had al-dawa? w bikam al-ʿilbe baʿd al-khasm?”
Bank “hawwalt alfein w khamsmiyye min hsaabi imbaareh w ma wislat.”
Telecom “al-baaqa khilsat w ana lissa bi-nuss al-shahr, shu al-hall?”
Bookings “biddi aʾajjel al-hajz min al-khamees la-l-sabt, nafs al-ghurfe.”

Step two: record them in three voices.

The same voice three times is not enough. You need a man’s voice, a woman’s voice, and the voice of someone over fifty. Transcription systems degrade noticeably with older voices, and that is a group nobody tests.

Step three: compute the error on what matters only.

Do not count every word. Count only the critical words:

Category Why it is critical
Numbers Amounts, quantities, times
Names People, products, places
Negation “it did not arrive” against “it arrived”
Dates and days Thursday against Saturday

A system that misses 40% of words but gets every number and every negation right may be usable. A system that misses 20% but flips a negation is a catastrophe. Overall word error rate is a misleading metric for any real application.

Step four: test the four traps.

The trap Example What it exposes
Short negation “ma biddi” against “biddi” One letter flips the meaning
Compound number “alfein w khamsmiyye” Number fragmentation
Uncommon name “biddi ahki maʿ ustaz Muhannad” Arabic names outside the common list
Foreign word “el-appointment muʾajjal” Code switching mid sentence

Step five: repeat in six months.

Transcription systems improve fast, but not evenly across dialects. Keep your twenty recordings as a fixed set and rerun them. This becomes your own ruler, and it is more useful than any global leaderboard, because it measures your dialect and your customers.

Step six: measure the gender and age bias.

A dimension no Arabic leaderboard measures, and one I have watched sink entire projects:

The speaker What usually happens
Man, 25 to 45, common dialect Best results
Woman, same age Slight degradation
Over 60 Heavy degradation
Child Severe degradation
A regional accent within the dialect Unpredictable

If your product serves a government or health service, a high share of your audience is elderly. And measuring on your office colleagues’ voices does not represent them.


Why formal Arabic fails in speech to begin with

I want to explain something that gets misread constantly: even the “good” figure, 27.95%, is not good.

The comparison Approximate error rate
English in clean conditions Under 5%
Modern Standard Arabic, this model 27.95%
Gulf Arabic, this model 59.92%

Look at the first and second rows. The best case in Arabic is roughly five times worse than the ordinary case in English. And that is before dialects enter the picture at all.

I raise this because the Arabic conversation dwells heavily on the formal-versus-dialect gap and forgets the larger one: the gap between Arabic as a whole and English. The first is an internal gap we can close with local data. The second is a structural gap that needs investment of a different order.


The practical verdict

If you are building an Arabic voice assistant, do not make voice the only path. Route every financial or critical decision through a text confirmation. That is not weak design. It is engineering that respects a sixty percent error rate.

And do not believe a single number for “Arabic”. Ask any vendor for a per-dialect breakdown, and anyone who does not have one has not tested your product.

If your audience is Gulf or Egyptian, know that you are in the two hardest cases specifically, and budget for correction and human review from the start.


Next in the series

One month later the French lab Mistral released Mixtral 8x7B and announced the languages it masters: English, French, Italian, German and Spanish. Five languages. The next article is about the list Arabic was not on, and what it means to be excluded explicitly rather than by omission.


Sources

  • OpenAI, New models and developer products announced at DevDay, November 6, 2023: openai.com
  • Whisper release v20231106 on GitHub: github.com/openai/whisper/releases
  • Wang, Alhmoud and Alqurishi, Open Universal Arabic ASR Leaderboard, Interspeech 2025, August 2025: isca-archive.org
  • Model card openai/whisper-large-v3 on Hugging Face