Which Language Does the Model Think In Before Answering You in Arabic? The DeepSeek R1 Case

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: January 20, 2025 | Published: February 3, 2025

The numbers first

MetricResultWhat it measuresSource
XReasoning, thinking-language match rate46.3% on AIME and 42.3% on GPQA for the largest R1-distilled modelDoes the model think in the language it was asked to?arXiv 2505.22888
XReasoning, after forcing the thinking languageMatch rises to 97.9%Effect of forcing the thinking languageSame source
The price: accuracyFalls from 25.5% to 17.0%Accuracy against thinking languageSame source
XReasoning, languages covered11Breadth of language coverageSame source
Is Arabic among them?NoArabic coverage in the benchmarkSame source
DialectalArabicMMLU, translating dialect questions to English+5.6 pointsEffect of translation on comprehensionLREC 2026
DialectalArabicMMLU, translating them to MSA+0.3 pointsEffect of translation on comprehensionLREC 2026
HELM Arabic, this model’s successor DeepSeek v3.1Among the top ten open modelsIndependent Arabic performanceStanford CRFM
The trade-off nobody discusses
match
accuracy
no intervention
46.3%
25.5%
forced thinking language 97.9%
17.0%
the more it thinks in your language, the more it errs

That trade-off is the subject here. And more important still: Arabic is not among the eleven languages measured. We do not even know the size of our own version of the problem, we infer it from other languages.


Why this question matters specifically in Arabic

Reasoning models display a chain of thought before answering. Ask in Arabic and you will often see the chain start in English, or start in Arabic and slide into English at the first hard analytical step. The final answer comes back in Arabic, so everything looks fine.

What does that mean in practice? Your Arabic question is translated internally, solved in an English space, and the result is translated back. Every translation step loses something:

your question in Levantine
↓ invisible internal translation
an English representation
← tone and intent are lost here
↓ reasoning in English
an English answer
↓ translation into Arabic
fluent Arabic, structurally sound, possibly foreign to your context

The numerical evidence for this path exists, and it is among the strongest findings I have read: in DialectalArabicMMLU, translating a dialect question into English raised model performance by 5.6 points, while translating it into MSA raised it by 0.3. If these models reasoned in Arabic, MSA would be the bridge. The bridge is English.


Model card

FieldDetail
NameDeepSeek R1
LabDeepSeek
ReleasedJanuary 20, 2025
DistinctionThe first widely used open reasoning model to expose its chain of thought
ImpactTopped US App Store downloads within a week of launch
Axis under testThe language of reasoning, before the language of the answer

This model’s importance for Arabic is not its score, it is its transparency. When a model shows its chain of thought, you can see with your own eyes what every closed model hides: which language the thinking actually happens in.


The test: four items

Item 1, the chain’s default language. Ask a cultural question in Levantine and read the chain. Log: which language did it start in, at which step did it switch, and what share of lines are Arabic versus English? That ratio takes a minute to compute and says a great deal.

Item 2, forcing. Ask explicitly: “think in Arabic, step by step.” Then measure two things at once: did it comply, and did answer quality change. Based on the XReasoning results, expect compliance with reduced accuracy. The number that matters to you is the size of that drop on your own tasks.

Item 3, cultural question versus logical question. Run the experiment twice: once with a mathematical reasoning question, once with a Levantine social situation. My hypothesis, after years of review, is that thinking in English hurts the cultural question more than the logical one, because a cultural question loses in translation what arithmetic does not. That is a testable hypothesis, and I publish it as a hypothesis rather than a result.

Item 4, the double comparison. Ask the same question in Arabic and in English in two separate sessions and compare. If the English answer is richer and sharper, you have practically measured what researchers call the English preference.


Reading the trade-off critically

The XReasoning result looks discouraging: either it thinks in your language or it is accurate. I read it slightly differently.

The model does not lose accuracy because Arabic or Swahili are “languages less capable of thought.” That is linguistic nonsense. It loses accuracy because its reasoning training happened in English, so forcing it onto a path it never trained on degrades it. The problem is in the data, not the language.

Which leads to an important practical conclusion: the problem is solvable, but solving it requires Arabic reasoning data, meaning chains of thought written in Arabic about Arabic problems. And that is exactly the kind of data nobody is collecting today, because everyone collects questions and answers, not reasoning paths.

One more note for anyone building now: if your product needs accuracy, let the model think however it wants and take a clean Arabic answer from it. If your product is educational and displays the steps to the user, you are forced into Arabic in the chain, and you should budget for the accuracy drop and offset it with review.


The practical verdict

  1. Inspect the chain’s language before judging the answer. One minute reveals whether your model understands your language or translates it.
  2. Do not force the thinking language unless displaying it is part of your product. And if you do, price the cost.
  3. Always ask in both Arabic and English in internal testing. The gap is a free measure of how deep the Arabic goes.
  4. If you are building Arabic data, build reasoning chains, not just question and answer pairs. That is the real gap in Arabic data today.

Next in the series

A step back to measure the distance: Gemini 2.5 Pro against Claude 3.5 Sonnet, the 2024 and early 2025 generation, and one question: what actually changed in Arabic over two years, and where did nothing change at all?


A second result supporting the hypothesis

I proposed above that thinking in English hurts cultural questions more than logical ones. That is my hypothesis and I present it as such. But there is a published result pointing the same way that deserves mention.

The AL-QASIDA study, which measured nine models across eight dialects, found that understanding outpaces generation in dialectal Arabic, inverting the usual pattern in generative models where production runs ahead of comprehension.

Read that with the XReasoning result and a consistent picture appears:

Arabic input
→
understood well
(comprehension is strong)
thinking
→
happens mostly in English
Arabic output
→
comes out MSA, not dialect
(generation is weak)

The weak link is not receiving Arabic, it is returning it. Which explains an experience every Arabic-speaking user knows: the model understands your question perfectly, then answers in a way that does not feel like yours.

A protocol for logging the chain’s language

ItemWhat to log
Language of the first line in the chainArabic or English
The step where it switchedA number
Share of Arabic linesA percentage
Did English terms appear inside Arabic thinkingYes or no
Final answer qualityOut of 5
Did quality change when forced into ArabicDifference in points

Five samples are enough for a stable number. Keep the table with the model name and date, because this behavior changes between versions faster than scores do.

What this means for building Arabic data

If the problem lies in training data rather than in the language, the fix is clear and expensive: reasoning chains written in Arabic.

Data typeAvailable todayWhat is missing
Arabic question and answer pairsAbundantNothing
General Arabic textAbundantSpelling quality
Arabic dialoguesBeginning to appearDialects beyond Egyptian and Maghrebi
Arabic reasoning chainsAlmost nonexistentEverything
Examples of saying I do not know, in ArabicNonexistentEverything

The last two rows are the real gap. Any Arab institution wanting lasting impact in this field will find that building a thousand human-written Arabic reasoning chains is worth more than training another model on the same data everyone already has.


From the test notebook: how I read a chain of thought

When a model shows its reasoning chain, I do not read it looking for errors. I read it looking for language. This is what I log in every session:

What I logWhy it matters
Language of the first lineReveals the model’s default state
The line where it switchedUsually the first hard analytical step
Ratio of Arabic lines to totalOne number, comparable across models
Whether it returned to Arabic at the endMany models think in English and summarize in Arabic
Whether the answer changed when forced into ArabicThis is the cost of forcing, measured on your tasks

The item I consider most important: cultural versus logical questions

I run the same experiment twice, once with a logical question and once with a Levantine cultural question, then compare. This is my hypothesis, and I present it as a testable hypothesis rather than an established result: thinking in English hurts the cultural question more than the logical one.

The reason I favor: a mathematics problem translates without loss, while “why did the host insist three times” loses in translation exactly the context that carries the answer. When a model thinks in English about an Arab situation, it is thinking about a translation of the situation rather than the situation.

I write the hypothesis down with its falsification criterion: if I find a model that gives the same cultural answer at the same quality whether it thinks in Arabic or English, the hypothesis fails.

The real gap in Arabic data

If the root of the problem is data rather than language, here is a plain inventory of what exists and what is missing:

Data typeStatus
General Arabic textAbundant, with a spelling problem rather than a volume problem
Arabic question and answer pairsAbundant
Dialectal dialoguesBeginning to appear, Levantine the thinnest
Labeled audio recordingsRare and expensive
Reasoning chains written in ArabicAlmost nonexistent
Examples of saying “I do not know” in an Arab contextNonexistent

The last two rows are what I consider the real opportunity. Any Arab institution that builds a thousand human-written Arabic reasoning chains about Arab problems has added something to the field that nobody holds today, and that is worth far more than training another model on the same data everyone already has.

Sources

  • When Models Reason in Your Language, the XReasoning benchmark: arxiv.org/abs/2505.22888
  • DialectalArabicMMLU, effects of translation to English and MSA, LREC 2026: arxiv.org/abs/2510.27543
  • HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • DeepSeek page, release dates: en.wikipedia.org/wiki/DeepSeek_(chatbot)
  • AL-QASIDA, understanding against generation in dialect: arxiv.org/abs/2412.04193