The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 59 at high reasoning, up 3 points from 3.7 Flash | Composite general intelligence | Artificial Analysis, September 2026 |
| Tool use | 45%, up 12 points on the previous version | Multi-step agentic tasks | Artificial Analysis |
| Speed | About 300 output tokens per second | Latency | Artificial Analysis |
| Price | $0.75 / $3.75 per million tokens through year end | Cost | Artificial Analysis |
| Supported inputs | Text, image, video and speech | Modalities | Artificial Analysis |
| Any published Arabic speech number | None | Its performance on Arabic audio | Survey of public sources |
The model takes audio in, and there is not one published figure on how it handles Arabic audio. To size the gap, here is what the specialist measurements say:
| Specialist measurement | Number | Source |
|---|---|---|
| NADI 2025, best participating ASR system | 35.68 WER, 12.20 CER | ArabicNLP 2025 |
| NADI 2025, spoken dialect identification | 79.8% accuracy | ArabicNLP 2025 |
| NADI 2025, diacritic restoration for spoken dialects | 55 WER, 13 CER | ArabicNLP 2025 |
| Teams submitting valid results | 8 of 44 registered | ArabicNLP 2025 |
| AHELM, best audio language model (Gemini 2.5 Pro) | 0.803 mean win rate across 10 dimensions | Stanford CRFM, 2025 |
| AHELM, rank of the simple ASR plus LLM baseline | 6th of 14 systems | Stanford CRFM |
That is the most important number here. More than a third of spoken words come out wrong in the best system competing on a task designed for Arabic dialects. Meanwhile Arabic voice products are being shipped today on the assumption that transcription is a solved problem.
Why WER is not enough to begin with
Word error rate measures one thing: did the characters come out right. In Arabic that is an incomplete test, for two reasons.
First, not all errors weigh the same. If the model hears ba’dar where the speaker said ma ba’dar, that is one word in the arithmetic and a complete reversal of meaning. If it writes biddi ruh where the speaker said biddi aruh, that is an error in the arithmetic and no error at all in understanding.
Second, transcription is not comprehension. So the ARAB-SPEECH axis in ARAB-LENS splits into four stages, and I do not let myself judge a voice model before walking through all of them:
The two stages nobody measures are exactly the two that decide whether a voice product works.
Model card
| Field | Detail |
|---|---|
| Name | Gemini 3.8 Flash |
| Lab | |
| Released | September 2, 2026 |
| Context | The fourth Flash model in under four months |
| Context window | 1M tokens |
| Inputs | Text, image, video, speech |
| Output | Text |
| Axis under test | ARAB-SPEECH, all four stages |
The test
Stage A, speech to text. Record thirty seconds of Levantine speech containing the three things models reliably mangle: an Arabic proper name (ʿAbd al-Rahman Abu al-Kheir), a compound spoken number (alfein w khamsmiyye w sabʿa w sabʿin), and a code switch (bukra ʿindi meeting al-saʿa tisʿa). Compute WER yourself. The arithmetic is simple: wrong words divided by total words.
Stage B, speech to meaning. Play a clip that says: “wallah kint biddi aji mbarih, bas sar maʿi shaghle w ma ‘dirt.” Do not ask for a transcript. Ask directly: why did the speaker not come, and did he apologize? A model that transcribes but cannot answer the question did not understand, it copied.
Stage C, speech to dialect. Same clip, new question: what region is this speaker from, and how confident are you? Here is a rule I impose on myself before I impose it on the model: I do not penalize it for saying “Levantine” rather than “Syrian.” Dialect and nationality are different things, and many recordings do not support a finer call. The real error is overconfidence, not caution.
Stage D, speech to intent. Record one phrase, “ei, tabʿan,” in five deliveries: enthusiasm, sarcasm, irritation, hesitation, and polite compliance. Then ask each time: what is the speaker actually communicating? This is the hardest stage on the axis, and it is where a model that hears separates from a person who understands.
Reading the AHELM results with Arabic eyes
One AHELM result matters enormously for anyone building in Arabic: the simple baseline, a transcription system followed by a language model, placed sixth among fourteen integrated audio systems, and proved more robust to noise than most of them.
In practice that means the long road, transcribe then understand, is still competitive and sometimes sturdier. For an Arabic product operating in real conditions, a cafe, a street, a phone call, that is not a technical detail, it is an architectural decision. The integrated audio model is faster and demos beautifully. The separate transcription pipeline is far easier to audit, correct, and hand to a human reviewer.
One more finding worth carrying over: AHELM detected a statistically significant gender bias in some speech tasks for the top-ranked model. “Best on average” does not mean “fair in every case,” and that lesson transfers directly to dialects. A model that is excellent on average can be specifically poor with rural speech or with older speakers.
The practical verdict
- Do not buy the phrase “understands spoken Arabic” without testing it. Record ten minutes of your actual users, not a trained reader in a quiet room, and measure on that.
- Separate transcription from understanding in your architecture. The separate baseline proved more robust under noise, and you will need the transcript for human review regardless.
- Allow uncertainty in dialect. Have your system say “Levantine, medium confidence” rather than guessing “Syrian” with high confidence. Overconfidence costs you a user who feels mislabeled.
- Measure tone by hand. No off-the-shelf metric exists for spoken intent, so build five fixed audio samples of your own and run every new model against them.
Next in the series
Back to pure text, and to the axis most neglected in Arabic evaluation: morphology, syntax and diacritics. DeepSeek V4 Pro goes on the table, with one simple question: does the model know the difference between darasa and darrasa when the vowel marks are gone?
What the research community measures in Arabic speech
Arabic speech is a small research field relative to the size of the language, and that shows in the numbers themselves: NADI 2025 had forty four teams register and only eight submit valid results. Eight teams for the first shared task in multidialectal Arabic speech.
That figure says more about the field than about the teams. Labeled Arabic audio is rare and expensive, and entering this area requires what most research groups do not have: native speakers able to distinguish a pronunciation error from a regional variant.
The Arabic speech literature also repeatedly names code-switching between Arabic and English or French as among the hardest problems facing multidialectal Arabic recognition. That matters particularly for anyone building in the Gulf or the Maghreb, where code-switching is not an exception, it is ordinary speech.
What practitioners say about tone
In published Arabic practitioner comparisons, particular models are repeatedly described as having good comprehension but a very formal tone, suited to official content. That sounds like a matter of taste, and in voice products it is fatal: an assistant that answers in newsreader register during an everyday conversation reads as cold or condescending, even when every word is correct.
So my evaluation of any Arabic voice product carries an item with nothing to do with accuracy: does this voice sound like someone I know? WER and CER cannot measure it, and five listeners settle it in five minutes.
The measurement log I recommend
Do not settle for a single WER. Log six fields per recording:
| Field | Why |
|---|---|
| Environment (quiet, cafe, car, phone) | The gap between them is your real production number |
| Dialect and sub-region | Performance varies between urban and rural inside one dialect |
| Speaker gender and age band | AHELM detected a statistically significant gender bias in the top-ranked model |
| Presence of proper names or numbers | The most common breaking points |
| Presence of code-switching | A documented challenge in the literature |
| Whether the meaning was understood, not whether text was copied | The difference between a working product and a transcriber |
Most teams skip the third field, and the consequence is discovering after launch that their system fails more often with older speakers and with women, which is a discovery that arrives far too late.
From the speech notebook: what I actually log
My fixed sample for Arabic speech testing is thirty seconds containing six deliberate elements, each of which breaks a different class of system:
| What I put in the recording | Why | The expected error |
|---|---|---|
| A compound proper name: ʿAbd al-Rahman Abu al-Kheir | Compound names get segmented wrongly | ʿAbd al-Rahman becomes ʿAbdel Rahman as two wrong units |
| A spoken number: alfein w khamsmiyye w sabʿa w sabʿin | Arabic numbers are spoken in reverse order | It emits 2577, or reads it digit by digit |
| A code switch: bukra ʿindi meeting al-saʿa tisʿa | Mid-sentence language transition | It transliterates meeting into Arabic letters or drops it |
| A Bedouin qaf: gultillak | Outside urban representation | It outputs qultu lak or broken text |
| Fast assimilation: shu baddak taʿmal halla’ | Fast speech swallows short vowels | baddak disappears or flattens |
| Real background noise | Because your user is not in a studio | Error sometimes doubles |
And I log two numbers per recording, not one: word error rate, then meaning-changing error rate. The second is what actually matters: ten vowel errors are less costly than one error on the negative particle ma.
An example showing why WER alone misleads
Take this spoken sentence:
“wallah ma qultillo inno ma yiji.”
If the system transcribes it as “wallah ma qultu lahu innahu ma yiji”, the arithmetic registers errors and the meaning is intact. If it transcribes it as “wallah qultu lahu innahu yiji”, the arithmetic registers only two word errors, and the meaning has inverted completely: the sentence now says the opposite of what was said.
The first system is worse on WER and better in reality. That is precisely why I insist on the second metric, and I know of no published leaderboard that computes it.
Accent from audio: what I accept and what I reject
When I ask for a dialect label from a recording, I accept this: “Levantine, medium confidence, with a coastal feature in the qaf.” I reject this: “Syrian, from Damascus”, because the recording does not carry enough to justify that certainty.
The reason is not pedantry. The error here is social rather than technical: millions of people in the region speak a dialect other than that of the place they now live, and a system that decides from thirty seconds will misclassify a refugee or a migrant, and then build a product decision on it.
Sources
- Artificial Analysis, Gemini 3.8 Flash release: artificialanalysis.ai/articles/gemini-3-8-flash
- NADI 2025, first multidialectal Arabic speech processing shared task: arxiv.org/abs/2509.02038
- AHELM, holistic evaluation of audio-language models, Stanford CRFM: arxiv.org/pdf/2508.21376
- DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
- Arabic practitioner comparisons on model tone: kuwaitai.net and arabie.ai