Fluency Is Not Understanding: Qwen3 and the Best Arabic Score Among Open Models

Reading Time: 11 min
18
Ehab Saleh | techkahwa.net Model released: April 28, 2025 | Published: May 12, 2025

The numbers first

MetricResultWhat it measuresSource
HELM Arabic, best open-weights modelQwen3 235B, mean 0.786Seven native Arabic benchmarksStanford CRFM and Arabic.AI, December 2025
HELM Arabic, open models in the top tenFour, including Qwen3 235B and Qwen3-Next 80BPresence of open models at the topStanford CRFM
OALL leaderboard, pretrained modelsQwen2.5-72B in the leadNative Arabic benchmarksOALL v2
Training data36 trillion tokens across 119 languages and dialectsBreadth of language coverageWikipedia, citing Qwen
Open sizes0.6B to 235B, Apache 2.0 licenseSelf-hosting and fine-tuning optionsSame source
DialectalArabicMMLU, dialect average47.7% against 51.9% MSA and 62.8% EnglishDialect comprehensionLREC 2026
Same models, three languages, three results
English
62.8
MSA
51.9
Dialects
47.7
DialectalArabicMMLU across 19 open models

An open-source model under Apache 2.0, trained on 119 languages, leading the open field on the strongest independent Arabic leaderboard. That is a real achievement and deserves acknowledgment, especially because it means any Arab team can download, run and fine-tune it without permission or a subscription.

The question this article asks is not about the number, it is about its meaning: 0.786 at what, exactly?


The central hypothesis in ARAB-LENS

After years of reviewing Arabic model output, I score two separate axes rather than one:

surface fluency
does the sentence look like beautiful Arabic?
semantic and cultural accuracy
is it actually correct in context?

Separate them and you get four cells, not two:

CorrectWrong
FluentGenuinely strong modelThe most dangerous of all
Not fluentUnderstands, speaks awkwardlyWeak on both

The dangerous cell is the top right: fluent and wrong. Polished Arabic, sound structures, rich vocabulary, carrying an unintended meaning or an inverted social judgment. A non-Arab reader will not catch it, a multiple-choice benchmark will not flag it, and a judging model will not warn you because it usually rewards beautiful language.

That cell is exactly what large multilingual models produce more than anyone else, because their Arabic fluency is excellent while their cultural representation is thin. A model trained on 119 languages knows what Arabic looks like. The question is how much it knows about what Arabic means.


Model card

FieldDetail
NameQwen3, flagship variant 235B-A22B
LabAlibaba Cloud
ReleasedApril 28, 2025
LicenseApache 2.0 for open variants
Training data36 trillion tokens across 119 languages and dialects
FeatureSwitching between thinking and non-thinking modes
Axis under testFluency versus understanding

The test: how to expose “fluent and wrong”

Item 1, the double score. For every response, assign two separate scores out of five: one for fluency, one for semantic and cultural accuracy. Never merge them. Then compute the share of responses scoring 4 or 5 on fluency and 2 or lower on accuracy. That share is your fluent-and-wrong rate, and to me it is the single most important number in evaluating any Arabic model.

Item 2, inverted interpretation. Give it a socially ambiguous situation, such as a host insisting a guest eat more, and ask for an interpretation. Score whether the interpretation is culturally right, entirely separately from how well it is written.

Item 3, the source question. After any cultural answer, ask: where did you learn this, and how confident are you? A fluent-and-wrong model keeps inventing confidently. An honest one separates what it has seen from what it is inferring.

Item 4, back-translation. Take its Arabic reply and ask for a literal English translation. Sometimes the translation reveals that the beautiful Arabic sentence was an English sentence translated in its head.

Item 5, cross-language comparison. Ask the same cultural question in Arabic and then in English, and compare. If the English answer is sharper and richer, you are looking at a model that thinks in English and writes in Arabic.

That last item is not my speculation. The DialectalArabicMMLU numbers support it directly: translating a dialect question into English raised performance by 5.6 points, while translating it into MSA raised it by 0.3. The road to a model understanding your dialect runs through English, not through MSA. That is the clearest numerical picture of what I call fluency without rootedness.


What deserves credit and what deserves caution

Credit: an open-weights model leading an independent Arabic leaderboard matters more for the region than it appears. Closed models cannot be audited, cannot be hosted locally, and cannot be tuned to a specific dialect. Four open models in the HELM Arabic top ten means the infrastructure for Arabic AI is now available to anyone who wants to build it.

Caution: the seven HELM Arabic benchmarks, good as they are, measure knowledge, grammar, safety, and grounded answering. Not one of them measures dialect generation, pragmatics, or conversation. So the 0.786 mean is entirely correct within its scope, and entirely silent outside it.


The practical verdict

  1. Never assign a single score to Arabic output. Always two: fluency and accuracy. The gap between them is your diagnosis.
  2. Ask the same question in Arabic and English. The gap tells you whether the model understands your language or translates it.
  3. An open model gives you a strategic option. If you hold clean Arabic data, fine-tuning an open model on your dialect is cheaper and more accurate than waiting for a global model to care about it.
  4. Be suspicious of beautiful text. In my reviews, the more elegant the Arabic, the more checking it needs, not less.

Next in the series

Llama 4 from Meta and the MSA leakage axis: how a model can follow your dialect request literally while violating it structurally at the same time, and why no published benchmark catches this.


The human evidence for the “fluent and wrong” cell

I wrote above that the dangerous cell in the matrix is fluent and wrong, and that multiple choice cannot catch it. In fact somebody did catch it, with human evaluation, in a phrase that is almost a definition of the cell.

The AL-QASIDA study, which measured nine models across eight dialects, described model output as:

Responses that are fluent and adequate in meaning, yet mostly not in the requested dialect.

That is a human judgment, not an automatic one. It says plainly that the model succeeds on two dimensions and fails on the third, and that the failure shows up not in the quality of the language but in its fit. The study added that understanding outpaces generation, inverting the usual pattern in generative models.

In numbers: most ADI2 scores fell below 50%, and GPT-4o did not clear 13% on half the dialects in the cross-lingual setting.

Double scoring, applied to three responses

The whole idea is to assign two scores rather than one. Here are three written examples illustrating the four cells, all answering the same question: “A host insisted three times that his guest eat, although the guest said he was full. What is happening?”

Response one: “This interaction indicates a violation of the guest’s personal boundaries, and the guest would be well advised to express his refusal clearly and firmly so the host respects his choice.” Fluency 5, the Arabic is clean and polished. Accuracy 1, the reading is socially inverted: the insistence is an obligation of hospitality rather than a violation, and the firm refusal is the insult. Cell: fluent and wrong, the most dangerous one.

Response two: “The insistence is part of the hospitality ritual, and the first refusal is an expected courtesy. The usual resolution is that the guest takes a little, praises the food, and blesses the host.” Fluency 5, accuracy 5. Cell: genuinely strong.

Response three: “The host want the guest eat more, this is Arab custom known, the guest say no but maybe he is hungry, he should take little.” Fluency 2, the structure is broken and agreement is unstable. Accuracy 4, the social understanding is essentially right. Cell: understands, speaks awkwardly.

This simple exercise is what I want every Arabic team to put into its review process: two scores per response, never one. Merging them hides exactly what you need to see.

How the cell and the rate are computed

fluency 4 or 5 + accuracy 4 or 5
→
genuinely strong
fluency 4 or 5 + accuracy 2 or less →
fluent and wrong
← this rate is your number
fluency 3 or less + accuracy 4 or 5 →
understands, speaks awkwardly
fluency 3 or less + accuracy 2 or less →
weak on both

A worked arithmetic illustration: if you classify twenty responses and three land in the fluent-and-wrong cell, your rate is 15%. One in every seven responses will look correct to an inattentive reader and will not be. Any product manager grasps that number immediately, unlike a general average.

The fluent-and-wrong rate is the number I want every Arabic team to compute and publish, and the only one that tells you how often your user will believe something false because it was written in beautiful Arabic.

What the QIMMA leaderboard said about this family

In April 2026 a model from the Qwen family led the Arabic QIMMA leaderboard with a mean of 68.06, on a board holding more than 52,000 samples across fourteen benchmarks with 99% natively Arabic content.

And the same leaderboard recorded an observation that completes the picture: Arabic-specialized models lead on cultural and linguistic tasks, while multilingual models lead on code. Topping the board overall does not mean topping every cell, and your choice should follow the shape of your task rather than the head of the table.


From the test notebook: how I design a question that exposes empty fluency

A good question on this axis has one property: its beautiful answer and its correct answer are not the same thing. I use three types:

Type one, a situation with two readings, one Western and one Arab. A host insisting, a relative visiting unannounced, a manager asking about something personal. A model reading the situation through a personal-boundaries lens writes an elegant and wrong paragraph, which is exactly what I want to measure.

Type two, a question about the origin of a word or expression. For example: “explain the origin and history of the word shanklish.” Etymology invites fabrication, because it looks like knowledge and is hard for a non-specialist to verify. What is required is that the model separate what is documented, what is a guess, and what it does not know.

Type three, a request for dialect phrasing with a precise meaning. For example: “apologize to your friend for being late, in Levantine, without using the word sorry.” Fluency and accuracy are tested together here: a model may write beautiful Levantine using an apology form nobody uses in that situation, or use the right form in broken Arabic.

Comparing two languages on the same question

This item gives the clearest evidence of fluency without rootedness. Ask the same cultural question twice, in Arabic and in English, in separate sessions, and compare.

What I often observe: the English answer is richer, sharper and more detailed, while the Arabic answer is more beautiful and thinner in substance. That matches the DialectalArabicMMLU result exactly: translating into English raises performance by 5.6 points, translating into MSA by 0.3.

The practical conclusion: if that gap is large for you, your Arabic product is running at half the model’s capability, and the fix is not a better Arabic prompt. It is an architecture that routes the hard question through English internally and returns the answer in Arabic, with human review on the tone.

Sources

  • HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • Open Arabic LLM Leaderboard v2: huggingface.co/blog/leaderboard-arabic-v2
  • DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
  • Qwen page, release dates and training data: en.wikipedia.org/wiki/Qwen
  • AL-QASIDA, fluency against dialect fidelity: arxiv.org/abs/2412.04193
  • QIMMA Arabic leaderboard: huggingface.co/blog/tiiuae/qimma-arabic-leaderboard