It Beat GPT-3.5 on Style and Lost to It on Knowledge: The AceGPT Case

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Paper released: September 21, 2023 | Published: October 5, 2023

The numbers first

Metric Result What it measures Source
AceGPT-13B-chat on Arabic Vicuna-80 100.88% of GPT-3.5’s performance Arabic instruction following, judged by GPT-4 AceGPT paper, NAACL 2024
AceGPT-13B-chat on Arabic AlpacaEval 97.95% of GPT-3.5 Same measure, different set Same paper
Jais-13B-chat on Arabic Vicuna-80 75.40% The direct Arabic competitor Same paper
AceGPT-13B-base on Arabic MMLU 40.45% General knowledge Same paper
GPT-3.5 Turbo on Arabic MMLU 49.07% The gap: 8.6 points to GPT-3.5 Same paper
AceGPT-13B-base on Arabic EXAMs 36.63% School examinations Same paper
GPT-3.5 Turbo on Arabic EXAMs 45.63% The gap: 9 points Same paper
AceGPT-13B-chat on ACVA, clean set 74.62% Arab cultural and value alignment Same paper
GPT-3.5 Turbo on ACVA, clean set 79.03% An American model wins on culture Same paper
LLaMA2-7B on Arabic MMLU, for reference 29.47% Guessing level Same paper

Read the first row and the last together: an Arabic model at thirteen billion parameters beat GPT-3.5 on style, lost to it on knowledge by eight points, and lost to it on Arab culture by four and a half.


Three axes, not one

AceGPT-13B against GPT-3.5, as a percentage of GPT-3.5
Arabic instruction following
100.9% win
Arabic general knowledge
82.4% loss
Arab culture and values
94.4% loss

That chart is the central idea of this whole series: fluency, knowledge and culture are three different things, measured by three different instruments, and you can win on one while falling on the others.


Model card

Item Detail
Model AceGPT, at 7 and 13 billion parameters
Institutions KAUST, the SRIBD institute, and the Chinese University of Hong Kong Shenzhen
Paper date September 21, 2023
Venue NAACL 2024
Base Continued training on top of Llama 2
Method Further pretraining in Arabic, then tuning on native Arabic instructions, then reinforcement learning with a reward model attuned to local culture
Central idea Localisation, not translation

The ARAB-LENS reading: the paper that said the right sentence

This paper contains what I consider the single most important sentence in the whole literature of computational Arabic:

“Relying on translated data may lead to localization issues, potentially undermining the integrity and applicability of the models in native contexts.”

And a second, stronger one:

“It is not just desirable but necessary to localize large language models and tailor them to a specific cultural environment.”

And they did not stop at saying it. They built the method on it: continued pretraining in Arabic, then native Arabic instructions rather than translated ones, then a reinforcement learning stage with a reward model attuned to local culture and values. That is the first time I saw a team put culture inside the training loop rather than outside it.

The result? An enormous jump in style. From 75.40% for Jais to 100.88% for AceGPT. An open thirteen billion parameter model was now writing Arabic that the judge preferred over GPT-3.5’s Arabic.

Then the second and third columns say something else entirely.

On knowledge: 40.45 against 49.07. On school exams: 36.63 against 45.63. And that is perfectly logical, because continued training does not plant new knowledge at density. It reshapes what is already there. The base was Llama 2, whose Arabic is under 0.005%. You cannot draw from the well what was never poured into it.

And on culture, which is the real surprise: AceGPT scored 74.62 on the clean ACVA set, and GPT-3.5 scored 79.03. So a model built specifically to align with Arab values, with a reward model tuned to local culture, lost to a general American model on an Arab cultural test.

How? My view is that there are two reasons. The first is that GPT-3.5 saw an enormous volume of text about the Arab world, even if most of it was in English, and cultural knowledge travels across languages. The second, and the more important one, is that ACVA is itself a multiple-choice exam, and multiple choice rewards knowing rather than behaving. A model that knows condolences last three days answers correctly. A model that knows how to speak at a condolence visit is not being measured at all.

That, in my judgement, is the most valuable thing this paper taught me: AceGPT won on what is measured by human judgement and lost on what is measured by multiple choice. And the question everyone evaluating an Arabic model should ask is: which of those two does my product need?


From the test notebook: separating fluency from knowledge from culture

This is my sharpest instrument, and I use it as a three-way ruler. Ask three questions about the same subject, each one measuring a different axis.

Subject: a condolence visit.

Axis The question What it exposes
Fluency “Write me a short condolence message in Levantine” Is the language natural?
Knowledge “How many days does the condolence period usually last, and what happens on the third day?” Does it know the fact?
Culture “My friend’s uncle died and I never met him. Do I have to go? And if I go, what exactly do I say when I first walk in?” Does it know how to behave?

A model that passes the first and fails the third is the most dangerous kind, because it is convincing and wrong.

Reference answers for the third: yes, you go, because the visit is a social duty towards your friend, not towards a deceased person you never met. On entering you say ʿazzam Allahu ajrakum or al-baqiyya bi-hayaatkum. You do not say congratulations, you do not ask for details about the cause of death, and you do not stay long if the room is full.

Second subject: an invitation.

Axis The question Reference answer
Fluency “Write a lunch invitation for my neighbours” Warm, short language
Knowledge “What is the traditional banquet dish in Jordan?” Mansaf
Culture “I invited someone and they said in shaa Allah. Do I set a place for them?” Yes, you do. Here in shaa Allah is a polite acceptance rather than a firm commitment, and the custom is to prepare and leave them an exit

Third subject: work.

Axis The question Reference answer
Fluency “Write me a polite resignation letter” Correct formal language
Knowledge “What is a mukhaalasa when leaving a job?” A final settlement of entitlements
Culture “My manager said khalleena nshoof when I asked for a raise. What does that mean?” Usually a deferred and polite refusal, not a promise

That last question is the best Arabic pragmatics test I know. “Khalleena nshoof” in an Arab workplace usually means a wrapped “no”. A model that reads it as a positive possibility is reading the words and not the situation.

Item four: one fast indicator.

If you have no time, ask one question: “What is the difference between someone telling you taʿaal ʿanna and telling you laazim teeji ʿanna?”

The first is a social pleasantry requiring no action. The second is a real invitation. A model that can tell them apart understands Arabic as it is lived, not as it is written.

Item five: the cultural reward test.

AceGPT used a reward model attuned to local culture. Test that effect directly with questions where a “globally correct” answer collides with a “locally appropriate” one:

The situation The global answer The locally appropriate one
“A colleague asked me for a loan and I do not want to” “Refuse clearly” A wrapped apology that preserves face
“My neighbour plays loud music at night” “Speak to him directly” An intermediary or a hint first
“My family want to know my salary” “Your privacy is your right” An answer that respects the family context
“My manager made a mistake in front of the team” “Correct him openly” Correct him privately

There is no single right answer here, and that is the point. A good model offers both options and explains the difference, rather than imposing one cultural template.

Item six: the dialect summarisation test.

Give the model a Levantine passage and ask for a Levantine summary. Most models summarise in formal Arabic even when asked otherwise, because summarisation is a task they learned in formal Arabic. Record the result: it is one of the clearest indicators of how deep the localisation goes.


The three-way ruler, applied to the table

Take the numbers at the top of the article and reorder them by the three axes:

Axis AceGPT-13B GPT-3.5 Winner
Fluency, Vicuna-80 100.88% of GPT-3.5 Reference AceGPT
Fluency, AlpacaEval 97.95% Reference Close
Knowledge, Arabic MMLU 40.45 49.07 GPT-3.5
Knowledge, EXAMs 36.63 45.63 GPT-3.5
Culture, clean ACVA 74.62 79.03 GPT-3.5

Five measurements. The Arabic model won one, drew one and lost three. But the one it won is the only one judged by humans, and the other four are multiple choice.

I will leave this question open for you: which of those two is closer to your real user’s experience, a human judging a written answer, or a choice among four options?


The practical verdict

If your task is writing, a model like AceGPT was the right choice in its time despite its size. Style improves quickly and cheaply with continued training.

If your task is answering questions or extracting information, scale and knowledge win, and no amount of linguistic tuning compensates.

And the rule I take away from this paper: before you choose a model, write on a piece of paper which of the three axes you actually need, fluency or knowledge or culture. Most of the failed projects I have seen chose a model that was excellent on an axis they did not need.


Next in the series

In November 2023 OpenAI released a new version of Whisper. Two years later a Saudi team measured it across five Arabic varieties and came back with two numbers: 27.95% error on Modern Standard Arabic and 59.92% on Gulf. The next article is about exactly that gap, and what it means for any Arabic voice application.


Sources

  • Huang et al., AceGPT: Localizing Large Language Models in Arabic, NAACL 2024: aclanthology.org/2024.naacl-long.450
  • First version on arXiv:2309.12053, September 21, 2023
  • Koto et al., ArabicMMLU, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334