The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| AceGPT-13B-chat on Arabic Vicuna-80 | 100.88% of GPT-3.5’s performance | Arabic instruction following, judged by GPT-4 | AceGPT paper, NAACL 2024 |
| AceGPT-13B-chat on Arabic AlpacaEval | 97.95% of GPT-3.5 | Same measure, different set | Same paper |
| Jais-13B-chat on Arabic Vicuna-80 | 75.40% | The direct Arabic competitor | Same paper |
| AceGPT-13B-base on Arabic MMLU | 40.45% | General knowledge | Same paper |
| GPT-3.5 Turbo on Arabic MMLU | 49.07% | The gap: 8.6 points to GPT-3.5 | Same paper |
| AceGPT-13B-base on Arabic EXAMs | 36.63% | School examinations | Same paper |
| GPT-3.5 Turbo on Arabic EXAMs | 45.63% | The gap: 9 points | Same paper |
| AceGPT-13B-chat on ACVA, clean set | 74.62% | Arab cultural and value alignment | Same paper |
| GPT-3.5 Turbo on ACVA, clean set | 79.03% | An American model wins on culture | Same paper |
| LLaMA2-7B on Arabic MMLU, for reference | 29.47% | Guessing level | Same paper |
Read the first row and the last together: an Arabic model at thirteen billion parameters beat GPT-3.5 on style, lost to it on knowledge by eight points, and lost to it on Arab culture by four and a half.
Three axes, not one
That chart is the central idea of this whole series: fluency, knowledge and culture are three different things, measured by three different instruments, and you can win on one while falling on the others.
Model card
| Item | Detail |
|---|---|
| Model | AceGPT, at 7 and 13 billion parameters |
| Institutions | KAUST, the SRIBD institute, and the Chinese University of Hong Kong Shenzhen |
| Paper date | September 21, 2023 |
| Venue | NAACL 2024 |
| Base | Continued training on top of Llama 2 |
| Method | Further pretraining in Arabic, then tuning on native Arabic instructions, then reinforcement learning with a reward model attuned to local culture |
| Central idea | Localisation, not translation |
The ARAB-LENS reading: the paper that said the right sentence
This paper contains what I consider the single most important sentence in the whole literature of computational Arabic:
“Relying on translated data may lead to localization issues, potentially undermining the integrity and applicability of the models in native contexts.”
And a second, stronger one:
“It is not just desirable but necessary to localize large language models and tailor them to a specific cultural environment.”
And they did not stop at saying it. They built the method on it: continued pretraining in Arabic, then native Arabic instructions rather than translated ones, then a reinforcement learning stage with a reward model attuned to local culture and values. That is the first time I saw a team put culture inside the training loop rather than outside it.
The result? An enormous jump in style. From 75.40% for Jais to 100.88% for AceGPT. An open thirteen billion parameter model was now writing Arabic that the judge preferred over GPT-3.5’s Arabic.
Then the second and third columns say something else entirely.
On knowledge: 40.45 against 49.07. On school exams: 36.63 against 45.63. And that is perfectly logical, because continued training does not plant new knowledge at density. It reshapes what is already there. The base was Llama 2, whose Arabic is under 0.005%. You cannot draw from the well what was never poured into it.
And on culture, which is the real surprise: AceGPT scored 74.62 on the clean ACVA set, and GPT-3.5 scored 79.03. So a model built specifically to align with Arab values, with a reward model tuned to local culture, lost to a general American model on an Arab cultural test.
How? My view is that there are two reasons. The first is that GPT-3.5 saw an enormous volume of text about the Arab world, even if most of it was in English, and cultural knowledge travels across languages. The second, and the more important one, is that ACVA is itself a multiple-choice exam, and multiple choice rewards knowing rather than behaving. A model that knows condolences last three days answers correctly. A model that knows how to speak at a condolence visit is not being measured at all.
That, in my judgement, is the most valuable thing this paper taught me: AceGPT won on what is measured by human judgement and lost on what is measured by multiple choice. And the question everyone evaluating an Arabic model should ask is: which of those two does my product need?
From the test notebook: separating fluency from knowledge from culture
This is my sharpest instrument, and I use it as a three-way ruler. Ask three questions about the same subject, each one measuring a different axis.
Subject: a condolence visit.
| Axis | The question | What it exposes |
|---|---|---|
| Fluency | “Write me a short condolence message in Levantine” | Is the language natural? |
| Knowledge | “How many days does the condolence period usually last, and what happens on the third day?” | Does it know the fact? |
| Culture | “My friend’s uncle died and I never met him. Do I have to go? And if I go, what exactly do I say when I first walk in?” | Does it know how to behave? |
A model that passes the first and fails the third is the most dangerous kind, because it is convincing and wrong.
Reference answers for the third: yes, you go, because the visit is a social duty towards your friend, not towards a deceased person you never met. On entering you say ʿazzam Allahu ajrakum or al-baqiyya bi-hayaatkum. You do not say congratulations, you do not ask for details about the cause of death, and you do not stay long if the room is full.
Second subject: an invitation.
| Axis | The question | Reference answer |
|---|---|---|
| Fluency | “Write a lunch invitation for my neighbours” | Warm, short language |
| Knowledge | “What is the traditional banquet dish in Jordan?” | Mansaf |
| Culture | “I invited someone and they said in shaa Allah. Do I set a place for them?” | Yes, you do. Here in shaa Allah is a polite acceptance rather than a firm commitment, and the custom is to prepare and leave them an exit |
Third subject: work.
| Axis | The question | Reference answer |
|---|---|---|
| Fluency | “Write me a polite resignation letter” | Correct formal language |
| Knowledge | “What is a mukhaalasa when leaving a job?” | A final settlement of entitlements |
| Culture | “My manager said khalleena nshoof when I asked for a raise. What does that mean?” | Usually a deferred and polite refusal, not a promise |
That last question is the best Arabic pragmatics test I know. “Khalleena nshoof” in an Arab workplace usually means a wrapped “no”. A model that reads it as a positive possibility is reading the words and not the situation.
Item four: one fast indicator.
If you have no time, ask one question: “What is the difference between someone telling you taʿaal ʿanna and telling you laazim teeji ʿanna?”
The first is a social pleasantry requiring no action. The second is a real invitation. A model that can tell them apart understands Arabic as it is lived, not as it is written.
Item five: the cultural reward test.
AceGPT used a reward model attuned to local culture. Test that effect directly with questions where a “globally correct” answer collides with a “locally appropriate” one:
| The situation | The global answer | The locally appropriate one |
|---|---|---|
| “A colleague asked me for a loan and I do not want to” | “Refuse clearly” | A wrapped apology that preserves face |
| “My neighbour plays loud music at night” | “Speak to him directly” | An intermediary or a hint first |
| “My family want to know my salary” | “Your privacy is your right” | An answer that respects the family context |
| “My manager made a mistake in front of the team” | “Correct him openly” | Correct him privately |
There is no single right answer here, and that is the point. A good model offers both options and explains the difference, rather than imposing one cultural template.
Item six: the dialect summarisation test.
Give the model a Levantine passage and ask for a Levantine summary. Most models summarise in formal Arabic even when asked otherwise, because summarisation is a task they learned in formal Arabic. Record the result: it is one of the clearest indicators of how deep the localisation goes.
The three-way ruler, applied to the table
Take the numbers at the top of the article and reorder them by the three axes:
| Axis | AceGPT-13B | GPT-3.5 | Winner |
|---|---|---|---|
| Fluency, Vicuna-80 | 100.88% of GPT-3.5 | Reference | AceGPT |
| Fluency, AlpacaEval | 97.95% | Reference | Close |
| Knowledge, Arabic MMLU | 40.45 | 49.07 | GPT-3.5 |
| Knowledge, EXAMs | 36.63 | 45.63 | GPT-3.5 |
| Culture, clean ACVA | 74.62 | 79.03 | GPT-3.5 |
Five measurements. The Arabic model won one, drew one and lost three. But the one it won is the only one judged by humans, and the other four are multiple choice.
I will leave this question open for you: which of those two is closer to your real user’s experience, a human judging a written answer, or a choice among four options?
The practical verdict
If your task is writing, a model like AceGPT was the right choice in its time despite its size. Style improves quickly and cheaply with continued training.
If your task is answering questions or extracting information, scale and knowledge win, and no amount of linguistic tuning compensates.
And the rule I take away from this paper: before you choose a model, write on a piece of paper which of the three axes you actually need, fluency or knowledge or culture. Most of the failed projects I have seen chose a model that was excellent on an axis they did not need.
Next in the series
In November 2023 OpenAI released a new version of Whisper. Two years later a Saudi team measured it across five Arabic varieties and came back with two numbers: 27.95% error on Modern Standard Arabic and 59.92% on Gulf. The next article is about exactly that gap, and what it means for any Arabic voice application.
Sources
- Huang et al., AceGPT: Localizing Large Language Models in Arabic, NAACL 2024: aclanthology.org/2024.naacl-long.450
- First version on arXiv:2309.12053, September 21, 2023
- Koto et al., ArabicMMLU, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334