“Bukra ʿindi meeting”: Testing GPT-5.1 on the Arabic We Actually Write

Reading Time: 10 min
18
Ehab Saleh | techkahwa.net Model released: November 12, 2025 | Published: November 26, 2025

The numbers first

MetricResultWhat it measuresSource
HELM ArabicAmong the closed models evaluated on seven native Arabic benchmarksIndependent Arabic performanceStanford CRFM and Arabic.AI, December 2025
HELM Arabic leaderArabic.AI LLM-XArabic performance on seven native benchmarksStanford CRFM
HELM Arabic best open modelQwen3 235B, mean 0.786Best Arabic performance among open modelsStanford CRFM
Artificial Analysis IndexGemini 3 Pro was three points ahead at its releaseGeneral intelligenceArtificial Analysis, November 2025
Code-switching, independent evaluation of an Arabic model4.92 of 5, the highest category of allMixing Arabic and EnglisharXiv 2508.17378, August 2025
MSA in the same evaluation4.74Quality of Modern Standard ArabicSame source
Levantine in the same evaluation2.73Levantine dialect fidelitySame source
Difficulty ranking according to published numbers
Code-switching
4.92
MSA
4.74
Dialects (avg)
4.21
Levantine
2.73

That ranking is surprising at first: code-switching is the easiest thing these models do in Arabic, not the hardest. The reason makes sense once you think about it. Half the sentence is English, and English is what these models are built on. The Arabic half of a sentence like “bukra ʿindi meeting” is short, simple, and abundant in training data.

But the high score hides the real failure, which is not in comprehension. It is in what the model hands back to you.


The failure that a 4.92 does not show

Write this to any strong model and ask for a reply:

“khalas ʿimilt submit lil-application, bas al-HR talabu minni al-transcript w ana ma baʿrif kif bitallʿo min al-jamʿa.”

It will almost certainly understand the request perfectly. That is the 4.92. Now look at the reply itself, and you will find one of three patterns:

PatternWhat it looks likeWhy it is wrong
Forced MSA“You may submit a request to obtain your academic transcript from the Deanship of Admissions and Registration”It converted your personal voice into an administrative bulletin
Over-translationIt renders submit, transcript and HR into ArabicYou did not write them in Arabic because nobody says them in Arabic in this context
Over-EnglishIt replies with English words in places you never used themIt is imitating the shape rather than the logic

The third pattern is the most diagnostic, because Arabic code-switching is not random. We do not mix the two languages arbitrarily, we follow unwritten rules:

professional and technical terms
→
stay English
(meeting, report, submit)
verbs, connectives, pronouns
→
stay Arabic
(ʿindi, baʿatli, bas)
emotion, blessings, courtesy
→
always Arabic (Allah yaʿtik al-ʿafiya)
numbers and dates
→
varies by country and context

A model that switches in the wrong places produces something comprehensible and unnatural at once, and that is exactly what makes a user feel a reply is “robotic” without being able to say why.


Model card

FieldDetail
NameGPT-5.1
LabOpenAI
ReleasedNovember 12, 2025, with more variants on November 19
VariantsInstant, Thinking, Pro, Codex-Mini, Codex-Max
Vendor claimA warmer personality and eight customizable personalities
Axis under testARAB-LING with ARAB-GEN: code-switching and voice preservation

The customizable personalities deserve a specifically Arabic test, because the question nobody has asked is: do these personalities behave in Arabic the way they behave in English? A “friendly” persona calibrated on English norms can, in Arabic, produce an excess of courtesy that reads as forced, or the reverse, a dryness that reads as rudeness.


The test: six items

Item 1, comprehension. The sentence above. Did it catch the actual problem, that he does not know the university’s procedure, rather than that he needs the word transcript defined?

Item 2, voice preservation. Ask for a reply “in my own style.” A good model hands back a sentence resembling yours: spoken Arabic with the English terms in their proper places. Trap: replying in formal MSA.

Item 3, selective translation. Ask it: which words in my sentence should stay English and which should be translated, and why? An excellent diagnostic item, because it reveals whether the model understands the logic of switching or merely imitates its shape.

Item 4, the reverse direction. Give it a formal MSA paragraph and ask it to convert it into everyday office Arabic as an employee in Amman would write it on WhatsApp.

Item 5, context limits. Ask for the same reply once to a colleague and once to an external client. Code-switching is fine in the first and may be inappropriate in the second. A model that does not differentiate does not understand register.

Item 6, the Arabizi trap. Write to it in Latin characters: “b3rf ino sa3b bs 3am 7awel.” Does it understand? And does it reply in Arabic script or imitate the encoding? Millions of Arabic speakers write this way daily, and there is not one published benchmark measuring this ability.


A research note worth stating

In Arabic speech research, code-switching between Arabic and English or French is repeatedly named as one of the core challenges in multidialectal Arabic speech recognition. In other words, what is relatively easy in text becomes hard in audio, because the system must detect a language transition mid-sentence spoken with an Arabic accent.

Put that beside the NADI 2025 figure, where the best word error rate on multidialectal Arabic speech recognition was 35.68, and you see why ARAB-LENS separates “code-switching in text” from “code-switching in speech.” The first is nearly solved. The second is the first thing that breaks in voice applications.


The practical verdict

  1. Do not measure code-switching by comprehension, measure it by voice. Comprehension is nearly guaranteed. Voice is what the user loses.
  2. Block forced term translation in your system prompt. One line does it: keep technical terms exactly as the user wrote them.
  3. Test Arabizi if your audience is young. A large share of user messages arrive in Latin characters, and your model either reads them or loses those users.
  4. Test custom personalities in Arabic specifically. English calibration for warmth does not transfer intact into an Arabic context.

Next in the series

Claude Opus 4.5 and the adversarial set: questions designed specifically to expose a model that has memorized the form. Does it admit when it does not know? And what does it do when handed an Arab custom that does not exist?


What localizers observe about code-switching

In Arabic practitioner tests of models on translation and localization, the same pattern recurs: weak handling of code-switching, and translating technical terms that are normally left in English. Alongside it come plural agreement errors, confusion with undiacritized text, and literal rendering of idioms.

I read that list as the inverse of what leaderboards measure. Leaderboards measure knowledge and grammar. The practitioner notices tone and terminology. The distance between the two lists is exactly what the user pays for.

The switching rules the model does not know

Arabic code-switching is not random. It follows unwritten rules that vary by domain. Here is a table you can place directly into your system prompt:

DomainStays EnglishStays Arabic
Technologydeploy, commit, bug, meeting, reportVerbs, connectives, pronouns
MedicineDrug names, test names and abbreviationsDescribing symptoms, feelings, asking how someone is
Educationassignment, quiz, GPA, semesterAssessment, encouragement, guidance
BusinessKPI, deadline, pipeline, invoiceCourtesy, apology, negotiation
EverydayApp and brand namesEverything else

The general rule I give any team: professional terminology stays as the user said it, and emotion and courtesy stay Arabic always. A model translating Allah yaʿtik al-ʿafiya into English inside an Arabic reply is wrong, and a model translating “report” into Arabic in a message to a colleague is also wrong, even though both replies are linguistically correct.

Arabizi: the completely absent axis

Millions of users write Arabic in Latin characters and digits: 3 for ʿayn, 7 for ha, 2 for hamza. Not one published benchmark measures this ability, not OALL, not HELM Arabic, not QIMMA.

The test protocol is simple:

ItemExampleRequired
Comprehensionb3rf ino sa3b bs 3am 7awelUnderstand it and reply in Arabic script
ConversionConvert the above into Arabic scriptAccurate transfer with no loss
Voice preservationReply in the same styleNot a formal bulletin
No blind imitationAny message from the examples aboveDo not reply in Arabizi unless asked

If your audience is young, or specifically in the Gulf and the Levant, this is not an optional test: a meaningful share of support messages arrive in exactly this form.


From the test notebook: Arabizi, the part nobody measures

Millions of users write Arabic in Latin letters and digits, especially in the Levant and the Gulf. Not one published benchmark measures this ability, not OALL, not HELM Arabic, not QIMMA. Here is the common encoding I test with:

SymbolLetterExample
2hamzaso2al means su’al
3ʿayn3arabi means ʿarabi
5 or 7′kha5alas means khalas
6ta6ayeb means tayyib
7ha7abibi means habibi
8 or 3′ghayn or qaf8alat means ghalat
9sad9a7i7 means sahih

And these are four messages I send as they are:

One: b3rf ino sa3b bs 3am 7awel Required: understand baʿrif inno saʿb bas ʿam hawil and reply in Arabic script in a matching dialect, not in MSA.

Two: 5alas 3melt submit lal application Required: understand the code switch inside the Arabizi, and keep submit and application as they are.

Three: wen9ar el maw3ed? Required: understand wein sar al-mawʿid and ask for clarification if the encoding is ambiguous, rather than guessing.

Four: Convert this message into Arabic script without changing its dialect. Many models fail here: they transliterate correctly and then “correct” the dialect into MSA, producing a message nobody wrote.

The switching rules I verify

Arabic code-switching is not random, and its logic can be tested:

Stays EnglishAlways stays Arabic
Professional terms: meeting, report, deadlineVerbs, connectives, pronouns
Product and app namesEmotion, blessings, courtesy
Abbreviations: HR, KPI, CVInterrogatives and negation

And my best diagnostic item is not comprehension, it is this direct question:

Which words in my sentence should stay English and which should be translated, and why?

A model that understands the logic answers with a rule. A model imitating the surface gives you an arbitrary list or translates everything.

The error that costs you a user

Allah yaʿtik al-ʿafiya is not translated inside an Arabic reply. Rendering it as “may God give you strength” in an Arabic context makes the whole reply look translated from another language. Conversely, translating “report” into taqrir in a message between two technical colleagues makes the reply sound formal and foreign.

The rule I give any team: terminology stays as the user said it, emotion stays Arabic. One line in the system prompt resolves most of these cases.

Sources

  • HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • Independent UI-level evaluation of an Arabic model, code-switching category: arxiv.org/abs/2508.17378
  • NADI 2025, challenges in multidialectal Arabic speech: arxiv.org/abs/2509.02038
  • Artificial Analysis, November 2025 comparisons: artificialanalysis.ai
  • Practitioner test of Arabic translation models: localazy.com/blog/ai-8-llm-arabic-models-tested-to-translate