The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| XReasoning, thinking-language match rate | 46.3% on AIME and 42.3% on GPQA for the largest R1-distilled model | Does the model think in the language it was asked to? | arXiv 2505.22888 |
| XReasoning, after forcing the thinking language | Match rises to 97.9% | Effect of forcing the thinking language | Same source |
| The price: accuracy | Falls from 25.5% to 17.0% | Accuracy against thinking language | Same source |
| XReasoning, languages covered | 11 | Breadth of language coverage | Same source |
| Is Arabic among them? | No | Arabic coverage in the benchmark | Same source |
| DialectalArabicMMLU, translating dialect questions to English | +5.6 points | Effect of translation on comprehension | LREC 2026 |
| DialectalArabicMMLU, translating them to MSA | +0.3 points | Effect of translation on comprehension | LREC 2026 |
| HELM Arabic, this model’s successor DeepSeek v3.1 | Among the top ten open models | Independent Arabic performance | Stanford CRFM |
That trade-off is the subject here. And more important still: Arabic is not among the eleven languages measured. We do not even know the size of our own version of the problem, we infer it from other languages.
Why this question matters specifically in Arabic
Reasoning models display a chain of thought before answering. Ask in Arabic and you will often see the chain start in English, or start in Arabic and slide into English at the first hard analytical step. The final answer comes back in Arabic, so everything looks fine.
What does that mean in practice? Your Arabic question is translated internally, solved in an English space, and the result is translated back. Every translation step loses something:
The numerical evidence for this path exists, and it is among the strongest findings I have read: in DialectalArabicMMLU, translating a dialect question into English raised model performance by 5.6 points, while translating it into MSA raised it by 0.3. If these models reasoned in Arabic, MSA would be the bridge. The bridge is English.
Model card
| Field | Detail |
|---|---|
| Name | DeepSeek R1 |
| Lab | DeepSeek |
| Released | January 20, 2025 |
| Distinction | The first widely used open reasoning model to expose its chain of thought |
| Impact | Topped US App Store downloads within a week of launch |
| Axis under test | The language of reasoning, before the language of the answer |
This model’s importance for Arabic is not its score, it is its transparency. When a model shows its chain of thought, you can see with your own eyes what every closed model hides: which language the thinking actually happens in.
The test: four items
Item 1, the chain’s default language. Ask a cultural question in Levantine and read the chain. Log: which language did it start in, at which step did it switch, and what share of lines are Arabic versus English? That ratio takes a minute to compute and says a great deal.
Item 2, forcing. Ask explicitly: “think in Arabic, step by step.” Then measure two things at once: did it comply, and did answer quality change. Based on the XReasoning results, expect compliance with reduced accuracy. The number that matters to you is the size of that drop on your own tasks.
Item 3, cultural question versus logical question. Run the experiment twice: once with a mathematical reasoning question, once with a Levantine social situation. My hypothesis, after years of review, is that thinking in English hurts the cultural question more than the logical one, because a cultural question loses in translation what arithmetic does not. That is a testable hypothesis, and I publish it as a hypothesis rather than a result.
Item 4, the double comparison. Ask the same question in Arabic and in English in two separate sessions and compare. If the English answer is richer and sharper, you have practically measured what researchers call the English preference.
Reading the trade-off critically
The XReasoning result looks discouraging: either it thinks in your language or it is accurate. I read it slightly differently.
The model does not lose accuracy because Arabic or Swahili are “languages less capable of thought.” That is linguistic nonsense. It loses accuracy because its reasoning training happened in English, so forcing it onto a path it never trained on degrades it. The problem is in the data, not the language.
Which leads to an important practical conclusion: the problem is solvable, but solving it requires Arabic reasoning data, meaning chains of thought written in Arabic about Arabic problems. And that is exactly the kind of data nobody is collecting today, because everyone collects questions and answers, not reasoning paths.
One more note for anyone building now: if your product needs accuracy, let the model think however it wants and take a clean Arabic answer from it. If your product is educational and displays the steps to the user, you are forced into Arabic in the chain, and you should budget for the accuracy drop and offset it with review.
The practical verdict
- Inspect the chain’s language before judging the answer. One minute reveals whether your model understands your language or translates it.
- Do not force the thinking language unless displaying it is part of your product. And if you do, price the cost.
- Always ask in both Arabic and English in internal testing. The gap is a free measure of how deep the Arabic goes.
- If you are building Arabic data, build reasoning chains, not just question and answer pairs. That is the real gap in Arabic data today.
Next in the series
A step back to measure the distance: Gemini 2.5 Pro against Claude 3.5 Sonnet, the 2024 and early 2025 generation, and one question: what actually changed in Arabic over two years, and where did nothing change at all?
A second result supporting the hypothesis
I proposed above that thinking in English hurts cultural questions more than logical ones. That is my hypothesis and I present it as such. But there is a published result pointing the same way that deserves mention.
The AL-QASIDA study, which measured nine models across eight dialects, found that understanding outpaces generation in dialectal Arabic, inverting the usual pattern in generative models where production runs ahead of comprehension.
Read that with the XReasoning result and a consistent picture appears:
The weak link is not receiving Arabic, it is returning it. Which explains an experience every Arabic-speaking user knows: the model understands your question perfectly, then answers in a way that does not feel like yours.
A protocol for logging the chain’s language
| Item | What to log |
|---|---|
| Language of the first line in the chain | Arabic or English |
| The step where it switched | A number |
| Share of Arabic lines | A percentage |
| Did English terms appear inside Arabic thinking | Yes or no |
| Final answer quality | Out of 5 |
| Did quality change when forced into Arabic | Difference in points |
Five samples are enough for a stable number. Keep the table with the model name and date, because this behavior changes between versions faster than scores do.
What this means for building Arabic data
If the problem lies in training data rather than in the language, the fix is clear and expensive: reasoning chains written in Arabic.
| Data type | Available today | What is missing |
|---|---|---|
| Arabic question and answer pairs | Abundant | Nothing |
| General Arabic text | Abundant | Spelling quality |
| Arabic dialogues | Beginning to appear | Dialects beyond Egyptian and Maghrebi |
| Arabic reasoning chains | Almost nonexistent | Everything |
| Examples of saying I do not know, in Arabic | Nonexistent | Everything |
The last two rows are the real gap. Any Arab institution wanting lasting impact in this field will find that building a thousand human-written Arabic reasoning chains is worth more than training another model on the same data everyone already has.
From the test notebook: how I read a chain of thought
When a model shows its reasoning chain, I do not read it looking for errors. I read it looking for language. This is what I log in every session:
| What I log | Why it matters |
|---|---|
| Language of the first line | Reveals the model’s default state |
| The line where it switched | Usually the first hard analytical step |
| Ratio of Arabic lines to total | One number, comparable across models |
| Whether it returned to Arabic at the end | Many models think in English and summarize in Arabic |
| Whether the answer changed when forced into Arabic | This is the cost of forcing, measured on your tasks |
The item I consider most important: cultural versus logical questions
I run the same experiment twice, once with a logical question and once with a Levantine cultural question, then compare. This is my hypothesis, and I present it as a testable hypothesis rather than an established result: thinking in English hurts the cultural question more than the logical one.
The reason I favor: a mathematics problem translates without loss, while “why did the host insist three times” loses in translation exactly the context that carries the answer. When a model thinks in English about an Arab situation, it is thinking about a translation of the situation rather than the situation.
I write the hypothesis down with its falsification criterion: if I find a model that gives the same cultural answer at the same quality whether it thinks in Arabic or English, the hypothesis fails.
The real gap in Arabic data
If the root of the problem is data rather than language, here is a plain inventory of what exists and what is missing:
| Data type | Status |
|---|---|
| General Arabic text | Abundant, with a spelling problem rather than a volume problem |
| Arabic question and answer pairs | Abundant |
| Dialectal dialogues | Beginning to appear, Levantine the thinnest |
| Labeled audio recordings | Rare and expensive |
| Reasoning chains written in Arabic | Almost nonexistent |
| Examples of saying “I do not know” in an Arab context | Nonexistent |
The last two rows are what I consider the real opportunity. Any Arab institution that builds a thousand human-written Arabic reasoning chains about Arab problems has added something to the field that nobody holds today, and that is worth far more than training another model on the same data everyone already has.
Sources
- When Models Reason in Your Language, the XReasoning benchmark: arxiv.org/abs/2505.22888
- DialectalArabicMMLU, effects of translation to English and MSA, LREC 2026: arxiv.org/abs/2510.27543
- HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
- DeepSeek page, release dates: en.wikipedia.org/wiki/DeepSeek_(chatbot)
- AL-QASIDA, understanding against generation in dialect: arxiv.org/abs/2412.04193