The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Modern Standard Arabic, speech in | Supported | Whether it hears you | Official language table, seamless_communication repository |
| Modern Standard Arabic, speech out | Supported | Whether it speaks to you | Same table |
| Egyptian Arabic, speech in | Supported | Whether it hears your dialect | Same table |
| Egyptian Arabic, speech out | Not supported | Whether it speaks your dialect | Same table |
| Moroccan Arabic, speech in | Supported | Whether it hears Darija | Same table |
| Moroccan Arabic, speech out | Not supported | Whether it can pronounce Darija | Same table |
| Improvement in direct speech-to-text translation | +20% BLEU over the previous state of the art | The model’s achievement in general, not in Arabic | Paper abstract, arXiv:2308.11596 |
| Error reduction against Whisper-Large-V2 | 45% on average across 77 overlapping languages | A global average, not an Arabic number | Same paper |
| SeamlessM4T v2, average Arabic error rate | 38.16% across six test sets | Independent evaluation, fourth of sixteen models | Open Universal Arabic ASR Leaderboard, Interspeech 2025 |
| SeamlessM4T v2 on SADA, Modern Standard Arabic | 39.87% error | The baseline | Same source |
| SeamlessM4T v2 on SADA, Egyptian | 72.98% error | The real gap | Same source |
Look at the last two rows: from 39.87 to 72.98. The error nearly doubles when you move from formal Arabic to Egyptian. And three quarters of the words being wrong means a transcript that is useless for anything.
The support map, as Meta published it
And that is the entire headline: the dialect is heard, not spoken.
Model card
| Item | Detail |
|---|---|
| Model | SeamlessM4T |
| Lab | Meta |
| Release date | August 22, 2023 |
| Paper | arXiv:2308.11596 |
| Tasks | Speech recognition, text translation, speech to text, speech to speech |
| Arabic varieties supported | Three: Modern Standard, Egyptian, Moroccan |
| What Meta published about Arabic specifically | No separate figure I was able to confirm |
| Source of the Arabic numbers | Independent academic evaluation, two years later |
The ARAB-LENS reading: the half acknowledgement
I will start with the fair part, because this model has earned it.
Putting Egyptian and Moroccan in the official language table was a real step, and it was not normal in 2023. Most systems at the time wrote “Arabic” on one line and stopped. Meta wrote three lines and gave each a code: arb, ary and arz. That is an explicit admission that Arabic is not one language, and such admissions are rare.
Then the acknowledgement stopped halfway.
Look at the table again. Egyptian and Moroccan are supported inbound only. You can speak Egyptian and it will understand you and give you text. It cannot speak Egyptian back. The only Arabic speech output is Modern Standard.
What does that mean in practice? Picture a call: you speak Egyptian, and the person on the other end hears formal Arabic. You say “I want to book a room for two nights” in Cairo street Arabic, and they hear a voice delivering it in the register of a news bulletin. The system succeeded technically and failed humanly, because nobody talks like that.
And here is my view: this is less a technical shortfall than a data shortfall, and a data shortfall is a decision shortfall. Speech synthesis needs carefully labelled audio in one variety, and that did not exist, and nobody decided to create it. Because the return on investment in a synthetic Moroccan voice is not calculated the same way as the return on Spanish.
The third observation, and the one that matters most methodologically: every shiny number in the paper is global, not Arabic. “Plus twenty percent” is an average across dozens of languages. “Forty five percent error reduction against Whisper” is an average across seventy seven languages. I tried to extract a single Arabic figure from the paper’s tables and could not reach a value I trusted enough to print. So when you read a headline saying “Meta improves Arabic translation by twenty percent”, know that the headline was invented.
The real Arabic numbers arrived two years later, from a team at Elm in Saudi Arabia, at Interspeech 2025. And they are uncomfortable: 38.16% average error, and 72.98% on Egyptian. That last figure means three words in every four came out wrong.
From the test notebook: testing Arabic speech without a lab
These items need nothing but your phone microphone.
Item one: one sentence in four voices.
Record the same sentence with four realizations of qaf and count how often the transcript comes out right:
| Realization | As spoken | Expected transcript |
|---|---|---|
| With qaf | “qaal li qabl qaleel” | qaal li qabl qaleel |
| With hamza | “ʾaal li ʾabl ʾaleel” | qaal li qabl qaleel |
| With hard g | “gaal li gabl galeel” | qaal li qabl qaleel |
| With kaf | “kaal li kabl kaleel” | qaal li qabl qaleel |
A good system produces the same line four times. A weak one writes what it heard, or invents words. This test is harsh and fair at once, because all four are correct Arabic.
Item two: spoken numbers.
Numbers are where Arabic transcription fails most, because the phrasing differs radically across dialects:
| MSA | Levantine | Egyptian | Gulf |
|---|---|---|---|
| ithnaan wa ʿishroon | tnein w ʿishreen | itnein w ʿishreen | thnein w ʿishreen |
| al-saaʿa al-thaalitha wa al-nisf | al-saaʿa talaata w nuss | al-saaʿa talaata w nuss | al-saaʿa thalaath w nuss |
| miʾa wa khamsoon riyaalan | miyye w khamseen leera | meet gineih w khamseen | miya w khamseen riyaal |
Record ten amounts and appointment times in your own dialect and compute the error rate. In my experience this is the highest error category of any, and it is exactly the category that matters for a real application: a booking, an invoice, an appointment.
Item three: a foreign word inside an Arabic sentence.
This is how we actually talk:
“baʿatli el-link ʿal-WhatsApp w baʿdein ʿamalna meeting ʿal-Zoom.”
“el-deadline bukra, w ana lissa ma khallast el-presentation.”
“hajazitlak appointment ʿand el-doctor al-saaʿa arbaʿa.”
A weak system does one of two things: it writes the foreign word in mangled Arabic letters, or it drops it and continues. Record which one happens, because the first is fixable and the second is catastrophic.
Item four: silence and hesitation.
Real Arabic speech is full of yaʿni, eih, shu ismo, Allah ykhalleek. Record a sentence with three of these and ask for a transcript. Some systems delete them, which is useful. Some convert them into words that carry meaning, which corrupts the whole transcript. Try: “yaʿni ana, eih, shu ismo, biddi aʾajjel al-mawʿed.”
Item five: the Arabic names test.
Transcription systems are trained on name lists, and Arabic names are usually outside them. Record ten names and measure:
| Category | Examples | Why it is hard |
|---|---|---|
| Very common | Muhammad, Ahmad, Fatima | Usually succeeds |
| Middling | Muhannad, Razan, Bashar | Errors begin |
| Rare | Dirar, Naila, Suhaib | Fails often |
| Compound | Abd al-Rahman, Umm Kulthum | Split into two words |
| With honorifics | Abu Khalid, al-Hajj Samir | The honorific gets dropped |
The fourth category is more dangerous than it looks: a system that writes “Abd” and “al-Rahman” separately will break any search in a customer database.
Item six: the realistic noise test.
Record the same sentence in four environments: a quiet room, a street, a moving car, and a phone call. The figure that matters is not the quiet room number but the spread across the four. A good system degrades gently. A weak one collapses suddenly in the third.
What it would actually take to build real Arabic speech
After reading these numbers, here is the list of what is genuinely missing, and it is shorter than you think:
| The need | State today | Who could close it |
|---|---|---|
| Precisely labelled dialect recordings | Very scarce | Universities and broadcasters |
| Balanced distribution across regions | Absent | Regional cooperation |
| Recordings of older voices | Nearly nonexistent | Community projects |
| Transcripts matching speech, not tidied | Scarce | Paid human transcription |
| Dialectal speech synthesis | Not commercially available | Private investment |
Note the last column: nothing on this list requires a technical invention. All of it requires fieldwork and funding. Which in my view is both the good news and the bad news: the problem is solvable, and nobody has decided to solve it.
The practical verdict
If you are building an Arabic voice application, assume from the outset that your target dialect will transcribe at roughly double the error rate of formal Arabic. Design the interface around that: confirm in text, read amounts back, ask a clarifying question instead of assuming comprehension.
And do not build formal Arabic speech output for a user who speaks in dialect. That is worse than written text. A formal voice in a dialect context reads as cold, official, and sometimes as mockery.
If you are evaluating a vendor, ask for their number on a named dialect set, not on “Arabic”. Anyone who has only one number for all of Arabic has not measured your dialect at all.
Next in the series
Eight days after this release, the UAE announced Jais: the first large language model deliberately built for Arabic, on a hundred and sixteen billion Arabic tokens. It is a number that looks very large. The next article takes it apart and explains how fifty five billion became a hundred and sixteen.
Sources
- Meta, Introducing SeamlessM4T, August 2023: about.fb.com/news/2023/08/seamlessm4t-ai-translation-model
- SeamlessM4T: Massively Multilingual and Multimodal Machine Translation, arXiv:2308.11596, August 22, 2023
- Official language table, facebookresearch/seamless_communication repository, docs/m4t/README.md
- Wang, Alhmoud and Alqurishi, Open Universal Arabic ASR Leaderboard, Interspeech 2025, August 2025: isca-archive.org