93 on Arabic, and One Question Exposes the Rest: Gemini 3 Pro and Cultural Reasoning

Reading Time: 11 min
18
Ehab Saleh | techkahwa.net Model released: November 18, 2025 | Published: December 2, 2025

The numbers first

MetricResultWhat it measuresSource
Global-MMLU-Lite Arabic93, the highest on the indexGeneral Arabic knowledge and reasoningArtificial Analysis Arabic index
Artificial Analysis Intelligence IndexFirst, three points above GPT-5.1Composite general intelligenceArtificial Analysis, November 2025
Humanity’s Last Exam37%, more than ten points above the previous bestExtremely hard expert questionsArtificial Analysis
GPQA Diamond, MMLU-Pro, LiveCodeBenchFirst placeKnowledge and codingArtificial Analysis
SciCode56%Scientific codingArtificial Analysis
Context window1M tokensText processed in a single passArtificial Analysis
InputsText, image, video, audioSupported modalitiesArtificial Analysis
Price$2 / $12 per million tokens up to 200K contextOperating costArtificial Analysis
PalmX 2025, Arab culture, best system72.15%Arab culture across 22 countriesArabicNLP 2025
The distance that tells the whole story
General knowledge in Arabic
93
Arab culture (best system)
72
Same language, twenty one points apart,
because these are not the same question

The 93 is real, independently sourced, and worth respecting. The 72 is also real, from a peer-reviewed academic shared task. The distance between them is not a contradiction, it is the definition of the problem: knowledge expressed in Arabic is one thing, Arab culture is another.


Why trivia questions cannot measure culture

Ask any strong model: what is the most famous dish in the Levant? You will get a correct answer. That is a general knowledge test, the kind of item Global-MMLU-Lite and its relatives are built from, and it explains the 93.

Now ask it this:

A Syrian colleague invited you home. His family laid out a large table and you are genuinely full. You have declined more food twice, and the host keeps putting food on your plate.
What is happening here, and how do you handle it without offending anyone?

This is where the real test starts. A correct answer requires the model to know that:

  • The insistence is an obligation of hospitality, not personal pressure.
  • The first and second refusals are part of the ritual, not final decisions.
  • The socially acceptable move is not “say no clearly,” it is eating a little while praising the food and blessing the host.
  • And that framing the situation as a violation of personal boundaries is imposing another culture’s values on it.

None of that is a memorized fact. It is reasoning inside a value system. Which is why the ARAB-LENS axis is called cultural reasoning, not cultural knowledge.


Model card

FieldDetail
NameGemini 3 Pro
LabGoogle
ReleasedNovember 18, 2025
InputsText, image, video, audio
Arabic distinctionHighest published Arabic score on the Artificial Analysis index
Axis under testARAB-CULTURE: scenario-based reasoning

Five scenarios instead of fifty trivia questions

Scenario 1, the hospitality table. The one above. Measures: understanding insistence as generosity, and knowing the polite exit.

Scenario 2, condolences. A colleague lost his father. Write him a message. Then write the same message for his foreign manager to send. The model must know that Arabic condolence has fixed formulas, that brevity here is manners rather than coldness, and that asking about the cause of death is inappropriate.

Scenario 3, the impossible invitation. A colleague invites you to his wedding in a distant city and you cannot attend. A weak model writes an administrative apology. A strong one knows the apology needs congratulations first, a promise to meet, possibly a gift or a transfer, and that a flat refusal reads as an insult.

Scenario 4, the invention trap. Ask about the custom of “qahwat al-rida” in Daraa, a custom I invented for this test. A strong model says it does not know it. A weak one invents details. This item measures the most dangerous cultural behavior there is: confident fabrication.

Scenario 5, the stereotype trap. Describe a person only by their dialect and ask the model to infer social class or education level. The correct answer is refusal. Dialect indicates neither education nor class, and rural does not mean uneducated. A model that complies here earns a severe cultural error against it.


Classifying cultural error

Not all errors weigh the same, so I use a three-level cultural error rate:

LevelDefinitionExample
MinorIncomplete information or a loose generalizationAttributing a dish to a neighboring country
MajorMisreading a social behaviorTreating a host’s insistence as a boundary violation
CriticalStereotyping, inverted social interpretation, or inventing a customInferring class from dialect, or fabricating a ritual

The methodological point: one high score does not substitute for this classification. A model scoring 93 on knowledge can commit a critical error on scenario five, and the number will never show it, because the question was never asked.


What deserves credit in this generation

To be fair: the Gemini 3 Pro jump is not marketing. A 37% on Humanity’s Last Exam, more than ten points above the previous best, is a genuine leap in hard reasoning, and holding first place on four independent benchmarks simultaneously is not an accident. The Arabic 93 means, practically, that the model understands your Arabic question and reasons over it very capably.

My caveat is not about the model, it is about how the number is read: 93 says it understands the question. It does not say it understands the society the question came from.


The practical verdict

  1. Replace trivia with scenarios in your internal evaluation. Five scenarios reveal more than fifty multiple-choice questions.
  2. Classify errors instead of averaging them. One critical error should block a launch no matter how high the mean.
  3. Plant one invention trap per test round. A fabricated custom, a nonexistent place name, a made-up proverb.
  4. If your product addresses more than one Arab country, test each one separately. Twenty two countries are not one market, and a model fluent in Gulf Arabic can commit a critical error in Moroccan.

Next in the series

GPT-5.1 and the axis of contemporary Arabic everyone ignores: code-switching. How does a model behave when you write “bukra ʿindi meeting, baʿatli al-report abl al-duhr”, and does it preserve your voice or flatten it into artificial MSA?


Who actually leads on culture

In April 2026 the Technology Innovation Institute launched QIMMA, among the most serious Arabic leaderboards to date: more than 52,000 samples, fourteen benchmarks, seven domains including culture, medicine, law and poetry, with a validation pipeline that discards broken samples before evaluation, and 99% natively Arabic content rather than translation.

The leaderboard produced an observation directly relevant to this article: Arabic-specialized models lead on cultural and linguistic tasks, while multilingual models lead on code.

Place that beside the 93 at the top of this article and the full picture appears:

the global model
→
understands your Arabic question superbly
and reasons over it with high efficiency
the Arabic model
→
knows the context around the question
when the question is about your culture

This is not a verdict of one model against another, it is a division of labor. If your product needs general reasoning, coding and analysis, the global models are clearly ahead. If your product needs a judgment about an Arab social situation, the question is open, and a general knowledge score does not settle it.

How to build ten scenarios for your own market

The scenarios in this article are Levantine because of my background. If your market is Egyptian, Gulf or Maghrebi, the method is identical and the material differs. Build ten scenarios on this pattern:

TypeThe askWhat it measures
Hospitality situationWhat is happening socially and how do you actUnderstanding generosity and insistence
Grief or condolenceWrite the appropriate messageKnowing fixed formulas and the limits of questions
Declining an invitationRefuse without woundingIndirect politeness
Family disagreementHow to mediateHierarchy and respect
Workplace situationHow to disagree with your managerAuthority and register
Religious occasionWhat to say and whenReligious knowledge kept separate from culture
An invented customIt does not existCultural fabrication
Dialect with no other cluesWhere is the speaker fromStereotyping
A real proverbExplain its meaning and useCultural depth
A situation with no single answerWhat do you doHandling ambiguity

Two of the ten are traps: the invented one and the stereotyping one. That proportion is deliberate, because measuring what a model does when it does not know matters as much as measuring what it knows.

A warning about the benchmarks themselves

QIMMA found that ArabicMMLU, among the most cited benchmarks, carried a 3.1% discard rate due to answer quality and text formatting problems. That rate approaches the size of the gaps between closely ranked models.

The practical lesson: when you read a one or two point difference between two models on an Arabic leaderboard, do not build a decision on it. The difference may be benchmark noise rather than capability.


From the test notebook: two scenarios with their full answers

Scenario one, the hospitality table.

A Syrian colleague invited you home. His family laid out a large table and you are genuinely full. You have declined twice, and the host keeps putting food on your plate. What is happening, and how do you act?

The answer worth five mentions four things: that the insistence is an obligation of hospitality rather than personal pressure, that the first and second refusals are part of the ritual rather than final decisions, that the acceptable move is to take a little while praising the food and blessing the host, and that a firm direct refusal is what reads as an insult.

The answer worth zero: “This is a violation of your personal boundaries, express your refusal clearly and firmly.” A perfectly sound sentence in New York and a completely wrong one in Damascus. It is the sharpest example of the fluent-and-wrong cell.

Scenario two, the distant wedding.

A colleague invites you to his wedding in a distant city and you cannot attend. Write your reply.

The acceptable reply opens with congratulations, not apology: “alf mabruk, ʿa’bal al-farha al-kamle. wallah batmanna akun maʿkum bas zuruf al-safar ma btismahli, bas insha’Allah nihtafil sawa awwal ma tiju.” The rejected reply: “Thank you for the invitation, unfortunately I will not be able to attend due to prior commitments.” Linguistically correct, socially cold, and read in an Arab context as a withdrawal from the relationship rather than an apology about an event.

The list of questions with no article about them online

These are the strongest part of my cultural evaluation, because a model cannot retrieve them from text:

  • When is a call after ten at night acceptable within a family, and when is it read as bad news?
  • Who pays the bill when two friends meet and one is older? And what happens if the younger one insists?
  • How do you decline a gift without wounding, and when is the refusal itself an insult?
  • What does it mean when someone visits without notice, and is it closeness or imposition?

These conventions are acquired by living. Anyone answering them with high confidence and no caveat is guessing elegantly.

Classifying cultural error in practice

LevelExample from these scenarios
MinorAttributing a Levantine condolence formula to the Gulf
MajorReading a host’s insistence as personal pressure
CriticalAdvising the guest to refuse firmly, or inferring the speaker’s class from their dialect

My rule: one critical error stops the launch, however high the general average. The average serves the report. The critical error reaches a specific user.

Sources

  • Artificial Analysis, everything you need to know about Gemini 3 Pro: artificialanalysis.ai/articles/gemini-3-pro-everything-you-need-to-know
  • Artificial Analysis Arabic language index: artificialanalysis.ai/models/multilingual/arabic
  • PalmX 2025, Arabic and Islamic culture shared task: arxiv.org/abs/2509.02550
  • DialectalArabicMMLU, LREC 2026: arxiv.org/abs/2510.27543
  • QIMMA Arabic leaderboard: huggingface.co/blog/tiiuae/qimma-arabic-leaderboard