Hearing Is Not Understanding: Testing Arabic Speech on Gemini 3.8 Flash

Reading Time: 11 min
18
Ehab Saleh | techkahwa.net Model released: September 2, 2026 | Published: September 16, 2026

The numbers first

MetricResultWhat it measuresSource
Artificial Analysis Intelligence Index59 at high reasoning, up 3 points from 3.7 FlashComposite general intelligenceArtificial Analysis, September 2026
Tool use45%, up 12 points on the previous versionMulti-step agentic tasksArtificial Analysis
SpeedAbout 300 output tokens per secondLatencyArtificial Analysis
Price$0.75 / $3.75 per million tokens through year endCostArtificial Analysis
Supported inputsText, image, video and speechModalitiesArtificial Analysis
Any published Arabic speech numberNoneIts performance on Arabic audioSurvey of public sources

The model takes audio in, and there is not one published figure on how it handles Arabic audio. To size the gap, here is what the specialist measurements say:

Specialist measurementNumberSource
NADI 2025, best participating ASR system35.68 WER, 12.20 CERArabicNLP 2025
NADI 2025, spoken dialect identification79.8% accuracyArabicNLP 2025
NADI 2025, diacritic restoration for spoken dialects55 WER, 13 CERArabicNLP 2025
Teams submitting valid results8 of 44 registeredArabicNLP 2025
AHELM, best audio language model (Gemini 2.5 Pro)0.803 mean win rate across 10 dimensionsStanford CRFM, 2025
AHELM, rank of the simple ASR plus LLM baseline6th of 14 systemsStanford CRFM
What WER = 35.68 means in practice
correct
64
wrong
36
Out of every hundred spoken dialectal words:
And that is the best system in a competition
built specifically for Arabic dialects, not a
general commercial model

That is the most important number here. More than a third of spoken words come out wrong in the best system competing on a task designed for Arabic dialects. Meanwhile Arabic voice products are being shipped today on the assumption that transcription is a solved problem.


Why WER is not enough to begin with

Word error rate measures one thing: did the characters come out right. In Arabic that is an incomplete test, for two reasons.

First, not all errors weigh the same. If the model hears ba’dar where the speaker said ma ba’dar, that is one word in the arithmetic and a complete reversal of meaning. If it writes biddi ruh where the speaker said biddi aruh, that is an error in the arithmetic and no error at all in understanding.

Second, transcription is not comprehension. So the ARAB-SPEECH axis in ARAB-LENS splits into four stages, and I do not let myself judge a voice model before walking through all of them:

speech → text
did it hear the characters?
measured by WER and CER
speech → meaning
did it understand?
measured by nobody
speech → dialect
does it know where from?
partly measured by NADI
speech → intent
did it catch the tone?
measured by nobody

The two stages nobody measures are exactly the two that decide whether a voice product works.


Model card

FieldDetail
NameGemini 3.8 Flash
LabGoogle
ReleasedSeptember 2, 2026
ContextThe fourth Flash model in under four months
Context window1M tokens
InputsText, image, video, speech
OutputText
Axis under testARAB-SPEECH, all four stages

The test

Stage A, speech to text. Record thirty seconds of Levantine speech containing the three things models reliably mangle: an Arabic proper name (ʿAbd al-Rahman Abu al-Kheir), a compound spoken number (alfein w khamsmiyye w sabʿa w sabʿin), and a code switch (bukra ʿindi meeting al-saʿa tisʿa). Compute WER yourself. The arithmetic is simple: wrong words divided by total words.

Stage B, speech to meaning. Play a clip that says: “wallah kint biddi aji mbarih, bas sar maʿi shaghle w ma ‘dirt.” Do not ask for a transcript. Ask directly: why did the speaker not come, and did he apologize? A model that transcribes but cannot answer the question did not understand, it copied.

Stage C, speech to dialect. Same clip, new question: what region is this speaker from, and how confident are you? Here is a rule I impose on myself before I impose it on the model: I do not penalize it for saying “Levantine” rather than “Syrian.” Dialect and nationality are different things, and many recordings do not support a finer call. The real error is overconfidence, not caution.

Stage D, speech to intent. Record one phrase, “ei, tabʿan,” in five deliveries: enthusiasm, sarcasm, irritation, hesitation, and polite compliance. Then ask each time: what is the speaker actually communicating? This is the hardest stage on the axis, and it is where a model that hears separates from a person who understands.


Reading the AHELM results with Arabic eyes

One AHELM result matters enormously for anyone building in Arabic: the simple baseline, a transcription system followed by a language model, placed sixth among fourteen integrated audio systems, and proved more robust to noise than most of them.

In practice that means the long road, transcribe then understand, is still competitive and sometimes sturdier. For an Arabic product operating in real conditions, a cafe, a street, a phone call, that is not a technical detail, it is an architectural decision. The integrated audio model is faster and demos beautifully. The separate transcription pipeline is far easier to audit, correct, and hand to a human reviewer.

One more finding worth carrying over: AHELM detected a statistically significant gender bias in some speech tasks for the top-ranked model. “Best on average” does not mean “fair in every case,” and that lesson transfers directly to dialects. A model that is excellent on average can be specifically poor with rural speech or with older speakers.


The practical verdict

  1. Do not buy the phrase “understands spoken Arabic” without testing it. Record ten minutes of your actual users, not a trained reader in a quiet room, and measure on that.
  2. Separate transcription from understanding in your architecture. The separate baseline proved more robust under noise, and you will need the transcript for human review regardless.
  3. Allow uncertainty in dialect. Have your system say “Levantine, medium confidence” rather than guessing “Syrian” with high confidence. Overconfidence costs you a user who feels mislabeled.
  4. Measure tone by hand. No off-the-shelf metric exists for spoken intent, so build five fixed audio samples of your own and run every new model against them.

Next in the series

Back to pure text, and to the axis most neglected in Arabic evaluation: morphology, syntax and diacritics. DeepSeek V4 Pro goes on the table, with one simple question: does the model know the difference between darasa and darrasa when the vowel marks are gone?


What the research community measures in Arabic speech

Arabic speech is a small research field relative to the size of the language, and that shows in the numbers themselves: NADI 2025 had forty four teams register and only eight submit valid results. Eight teams for the first shared task in multidialectal Arabic speech.

That figure says more about the field than about the teams. Labeled Arabic audio is rare and expensive, and entering this area requires what most research groups do not have: native speakers able to distinguish a pronunciation error from a regional variant.

The Arabic speech literature also repeatedly names code-switching between Arabic and English or French as among the hardest problems facing multidialectal Arabic recognition. That matters particularly for anyone building in the Gulf or the Maghreb, where code-switching is not an exception, it is ordinary speech.

What practitioners say about tone

In published Arabic practitioner comparisons, particular models are repeatedly described as having good comprehension but a very formal tone, suited to official content. That sounds like a matter of taste, and in voice products it is fatal: an assistant that answers in newsreader register during an everyday conversation reads as cold or condescending, even when every word is correct.

So my evaluation of any Arabic voice product carries an item with nothing to do with accuracy: does this voice sound like someone I know? WER and CER cannot measure it, and five listeners settle it in five minutes.

The measurement log I recommend

Do not settle for a single WER. Log six fields per recording:

FieldWhy
Environment (quiet, cafe, car, phone)The gap between them is your real production number
Dialect and sub-regionPerformance varies between urban and rural inside one dialect
Speaker gender and age bandAHELM detected a statistically significant gender bias in the top-ranked model
Presence of proper names or numbersThe most common breaking points
Presence of code-switchingA documented challenge in the literature
Whether the meaning was understood, not whether text was copiedThe difference between a working product and a transcriber

Most teams skip the third field, and the consequence is discovering after launch that their system fails more often with older speakers and with women, which is a discovery that arrives far too late.


From the speech notebook: what I actually log

My fixed sample for Arabic speech testing is thirty seconds containing six deliberate elements, each of which breaks a different class of system:

What I put in the recordingWhyThe expected error
A compound proper name: ʿAbd al-Rahman Abu al-KheirCompound names get segmented wronglyʿAbd al-Rahman becomes ʿAbdel Rahman as two wrong units
A spoken number: alfein w khamsmiyye w sabʿa w sabʿinArabic numbers are spoken in reverse orderIt emits 2577, or reads it digit by digit
A code switch: bukra ʿindi meeting al-saʿa tisʿaMid-sentence language transitionIt transliterates meeting into Arabic letters or drops it
A Bedouin qaf: gultillakOutside urban representationIt outputs qultu lak or broken text
Fast assimilation: shu baddak taʿmal halla’Fast speech swallows short vowelsbaddak disappears or flattens
Real background noiseBecause your user is not in a studioError sometimes doubles

And I log two numbers per recording, not one: word error rate, then meaning-changing error rate. The second is what actually matters: ten vowel errors are less costly than one error on the negative particle ma.

An example showing why WER alone misleads

Take this spoken sentence:

“wallah ma qultillo inno ma yiji.”

If the system transcribes it as “wallah ma qultu lahu innahu ma yiji”, the arithmetic registers errors and the meaning is intact. If it transcribes it as “wallah qultu lahu innahu yiji”, the arithmetic registers only two word errors, and the meaning has inverted completely: the sentence now says the opposite of what was said.

The first system is worse on WER and better in reality. That is precisely why I insist on the second metric, and I know of no published leaderboard that computes it.

Accent from audio: what I accept and what I reject

When I ask for a dialect label from a recording, I accept this: “Levantine, medium confidence, with a coastal feature in the qaf.” I reject this: “Syrian, from Damascus”, because the recording does not carry enough to justify that certainty.

The reason is not pedantry. The error here is social rather than technical: millions of people in the region speak a dialect other than that of the place they now live, and a system that decides from thirty seconds will misclassify a refugee or a migrant, and then build a product decision on it.

Sources

  • Artificial Analysis, Gemini 3.8 Flash release: artificialanalysis.ai/articles/gemini-3-8-flash
  • NADI 2025, first multidialectal Arabic speech processing shared task: arxiv.org/abs/2509.02038
  • AHELM, holistic evaluation of audio-language models, Stanford CRFM: arxiv.org/pdf/2508.21376
  • DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
  • Arabic practitioner comparisons on model tone: kuwaitai.net and arabie.ai