The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| HELM Arabic | Among the closed models evaluated on seven native Arabic benchmarks | Independent Arabic performance | Stanford CRFM and Arabic.AI, December 2025 |
| HELM Arabic leader | Arabic.AI LLM-X | Arabic performance on seven native benchmarks | Stanford CRFM |
| HELM Arabic best open model | Qwen3 235B, mean 0.786 | Best Arabic performance among open models | Stanford CRFM |
| Artificial Analysis Index | Gemini 3 Pro was three points ahead at its release | General intelligence | Artificial Analysis, November 2025 |
| Code-switching, independent evaluation of an Arabic model | 4.92 of 5, the highest category of all | Mixing Arabic and English | arXiv 2508.17378, August 2025 |
| MSA in the same evaluation | 4.74 | Quality of Modern Standard Arabic | Same source |
| Levantine in the same evaluation | 2.73 | Levantine dialect fidelity | Same source |
That ranking is surprising at first: code-switching is the easiest thing these models do in Arabic, not the hardest. The reason makes sense once you think about it. Half the sentence is English, and English is what these models are built on. The Arabic half of a sentence like “bukra ʿindi meeting” is short, simple, and abundant in training data.
But the high score hides the real failure, which is not in comprehension. It is in what the model hands back to you.
The failure that a 4.92 does not show
Write this to any strong model and ask for a reply:
“khalas ʿimilt submit lil-application, bas al-HR talabu minni al-transcript w ana ma baʿrif kif bitallʿo min al-jamʿa.”
It will almost certainly understand the request perfectly. That is the 4.92. Now look at the reply itself, and you will find one of three patterns:
| Pattern | What it looks like | Why it is wrong |
|---|---|---|
| Forced MSA | “You may submit a request to obtain your academic transcript from the Deanship of Admissions and Registration” | It converted your personal voice into an administrative bulletin |
| Over-translation | It renders submit, transcript and HR into Arabic | You did not write them in Arabic because nobody says them in Arabic in this context |
| Over-English | It replies with English words in places you never used them | It is imitating the shape rather than the logic |
The third pattern is the most diagnostic, because Arabic code-switching is not random. We do not mix the two languages arbitrarily, we follow unwritten rules:
A model that switches in the wrong places produces something comprehensible and unnatural at once, and that is exactly what makes a user feel a reply is “robotic” without being able to say why.
Model card
| Field | Detail |
|---|---|
| Name | GPT-5.1 |
| Lab | OpenAI |
| Released | November 12, 2025, with more variants on November 19 |
| Variants | Instant, Thinking, Pro, Codex-Mini, Codex-Max |
| Vendor claim | A warmer personality and eight customizable personalities |
| Axis under test | ARAB-LING with ARAB-GEN: code-switching and voice preservation |
The customizable personalities deserve a specifically Arabic test, because the question nobody has asked is: do these personalities behave in Arabic the way they behave in English? A “friendly” persona calibrated on English norms can, in Arabic, produce an excess of courtesy that reads as forced, or the reverse, a dryness that reads as rudeness.
The test: six items
Item 1, comprehension. The sentence above. Did it catch the actual problem, that he does not know the university’s procedure, rather than that he needs the word transcript defined?
Item 2, voice preservation. Ask for a reply “in my own style.” A good model hands back a sentence resembling yours: spoken Arabic with the English terms in their proper places. Trap: replying in formal MSA.
Item 3, selective translation. Ask it: which words in my sentence should stay English and which should be translated, and why? An excellent diagnostic item, because it reveals whether the model understands the logic of switching or merely imitates its shape.
Item 4, the reverse direction. Give it a formal MSA paragraph and ask it to convert it into everyday office Arabic as an employee in Amman would write it on WhatsApp.
Item 5, context limits. Ask for the same reply once to a colleague and once to an external client. Code-switching is fine in the first and may be inappropriate in the second. A model that does not differentiate does not understand register.
Item 6, the Arabizi trap. Write to it in Latin characters: “b3rf ino sa3b bs 3am 7awel.” Does it understand? And does it reply in Arabic script or imitate the encoding? Millions of Arabic speakers write this way daily, and there is not one published benchmark measuring this ability.
A research note worth stating
In Arabic speech research, code-switching between Arabic and English or French is repeatedly named as one of the core challenges in multidialectal Arabic speech recognition. In other words, what is relatively easy in text becomes hard in audio, because the system must detect a language transition mid-sentence spoken with an Arabic accent.
Put that beside the NADI 2025 figure, where the best word error rate on multidialectal Arabic speech recognition was 35.68, and you see why ARAB-LENS separates “code-switching in text” from “code-switching in speech.” The first is nearly solved. The second is the first thing that breaks in voice applications.
The practical verdict
- Do not measure code-switching by comprehension, measure it by voice. Comprehension is nearly guaranteed. Voice is what the user loses.
- Block forced term translation in your system prompt. One line does it: keep technical terms exactly as the user wrote them.
- Test Arabizi if your audience is young. A large share of user messages arrive in Latin characters, and your model either reads them or loses those users.
- Test custom personalities in Arabic specifically. English calibration for warmth does not transfer intact into an Arabic context.
Next in the series
Claude Opus 4.5 and the adversarial set: questions designed specifically to expose a model that has memorized the form. Does it admit when it does not know? And what does it do when handed an Arab custom that does not exist?
What localizers observe about code-switching
In Arabic practitioner tests of models on translation and localization, the same pattern recurs: weak handling of code-switching, and translating technical terms that are normally left in English. Alongside it come plural agreement errors, confusion with undiacritized text, and literal rendering of idioms.
I read that list as the inverse of what leaderboards measure. Leaderboards measure knowledge and grammar. The practitioner notices tone and terminology. The distance between the two lists is exactly what the user pays for.
The switching rules the model does not know
Arabic code-switching is not random. It follows unwritten rules that vary by domain. Here is a table you can place directly into your system prompt:
| Domain | Stays English | Stays Arabic |
|---|---|---|
| Technology | deploy, commit, bug, meeting, report | Verbs, connectives, pronouns |
| Medicine | Drug names, test names and abbreviations | Describing symptoms, feelings, asking how someone is |
| Education | assignment, quiz, GPA, semester | Assessment, encouragement, guidance |
| Business | KPI, deadline, pipeline, invoice | Courtesy, apology, negotiation |
| Everyday | App and brand names | Everything else |
The general rule I give any team: professional terminology stays as the user said it, and emotion and courtesy stay Arabic always. A model translating Allah yaʿtik al-ʿafiya into English inside an Arabic reply is wrong, and a model translating “report” into Arabic in a message to a colleague is also wrong, even though both replies are linguistically correct.
Arabizi: the completely absent axis
Millions of users write Arabic in Latin characters and digits: 3 for ʿayn, 7 for ha, 2 for hamza. Not one published benchmark measures this ability, not OALL, not HELM Arabic, not QIMMA.
The test protocol is simple:
| Item | Example | Required |
|---|---|---|
| Comprehension | b3rf ino sa3b bs 3am 7awel | Understand it and reply in Arabic script |
| Conversion | Convert the above into Arabic script | Accurate transfer with no loss |
| Voice preservation | Reply in the same style | Not a formal bulletin |
| No blind imitation | Any message from the examples above | Do not reply in Arabizi unless asked |
If your audience is young, or specifically in the Gulf and the Levant, this is not an optional test: a meaningful share of support messages arrive in exactly this form.
From the test notebook: Arabizi, the part nobody measures
Millions of users write Arabic in Latin letters and digits, especially in the Levant and the Gulf. Not one published benchmark measures this ability, not OALL, not HELM Arabic, not QIMMA. Here is the common encoding I test with:
| Symbol | Letter | Example |
|---|---|---|
| 2 | hamza | so2al means su’al |
| 3 | ʿayn | 3arabi means ʿarabi |
| 5 or 7′ | kha | 5alas means khalas |
| 6 | ta | 6ayeb means tayyib |
| 7 | ha | 7abibi means habibi |
| 8 or 3′ | ghayn or qaf | 8alat means ghalat |
| 9 | sad | 9a7i7 means sahih |
And these are four messages I send as they are:
One: b3rf ino sa3b bs 3am 7awel Required: understand baʿrif inno saʿb bas ʿam hawil and reply in Arabic script in a matching dialect, not in MSA.
Two: 5alas 3melt submit lal application Required: understand the code switch inside the Arabizi, and keep submit and application as they are.
Three: wen9ar el maw3ed? Required: understand wein sar al-mawʿid and ask for clarification if the encoding is ambiguous, rather than guessing.
Four: Convert this message into Arabic script without changing its dialect. Many models fail here: they transliterate correctly and then “correct” the dialect into MSA, producing a message nobody wrote.
The switching rules I verify
Arabic code-switching is not random, and its logic can be tested:
| Stays English | Always stays Arabic |
|---|---|
| Professional terms: meeting, report, deadline | Verbs, connectives, pronouns |
| Product and app names | Emotion, blessings, courtesy |
| Abbreviations: HR, KPI, CV | Interrogatives and negation |
And my best diagnostic item is not comprehension, it is this direct question:
Which words in my sentence should stay English and which should be translated, and why?
A model that understands the logic answers with a rule. A model imitating the surface gives you an arbitrary list or translates everything.
The error that costs you a user
Allah yaʿtik al-ʿafiya is not translated inside an Arabic reply. Rendering it as “may God give you strength” in an Arabic context makes the whole reply look translated from another language. Conversely, translating “report” into taqrir in a message between two technical colleagues makes the reply sound formal and foreign.
The rule I give any team: terminology stays as the user said it, emotion stays Arabic. One line in the system prompt resolves most of these cases.
Sources
- HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
- Independent UI-level evaluation of an Arabic model, code-switching category: arxiv.org/abs/2508.17378
- NADI 2025, challenges in multidialectal Arabic speech: arxiv.org/abs/2509.02038
- Artificial Analysis, November 2025 comparisons: artificialanalysis.ai
- Practitioner test of Arabic translation models: localazy.com/blog/ai-8-llm-arabic-models-tested-to-translate