Arabic Is the First Name on the List: Aya 23 and the Bet on Depth Over Breadth

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Model released: May 23, 2024 | Published: June 6, 2024

The numbers first

Metric Result What it measures Source
Number of languages 23 The model’s scope Aya 23 technical report
Arabic’s position on the list First Alphabetical in English, but named explicitly Same report
The languages Arabic, Chinese, Czech, Dutch, English, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Korean, Persian, Polish, Portuguese, Romanian, Russian, Spanish, Turkish, Ukrainian, Vietnamese Who is in and who is out Same report
Its predecessor Aya 101 101 languages The comparison point Same report
Sizes 8 and 35 billion parameters Accessibility Same report
The declared thesis “An experiment in depth vs breadth, exploring the impact of allocating more capacity to fewer languages” The core of the paper, quoted verbatim Paper abstract
Release date May 23, 2024 Confirmed on arXiv and the Cohere blog arXiv:2405.15032
A separate Arabic figure I could confirm None I tried three times and could not establish one well enough to publish My own methodological note

That last row matters, and I will not hide it: a model that puts Arabic first on its list, and I could not confirm a separate Arabic number for it that I trust.


Depth against breadth

Model capacity spread across languages, a principle not a measurement
Aya 101
101 languages, a sliver each
Aya 23
23 languages, a larger share each

This chart illustrates the idea the paper argues, not a published measurement. And the idea is simple and strong: one language out of twenty three gets more than one language out of a hundred and one.


Model card

Item Detail
Model Aya 23, at 8 and 35 billion parameters
Lab Cohere For AI and Cohere
Release date May 23, 2024
Paper arXiv:2405.15032
Licence Open weights
Arabic Explicitly named among twenty three languages
What was published about Arabic Its presence on the list, which is confirmed
What I could not confirm Any separate Arabic performance figure

The ARAB-LENS reading: why a short list matters to you more than a long one

This article is about an idea more than a number, and the idea earns its place.

This model’s predecessor, Aya 101, covered a hundred and one languages. At first glance that looks like an enormous achievement. A hundred languages! Coverage of the whole world!

Then the same team comes back in under a year and says: we tried the opposite. Twenty three languages only, and more capacity for each. The abstract calls it, in its own words, “an experiment in depth vs breadth”.

Why does that matter to an Arabic speaker specifically?

Because in long lists Arabic is treated as a minority language, and in short lists it is treated as a principal one. The difference is not in the naming but in the allocation: more data, sharper evaluation, native reviewers, and a heavier weight in training decisions.

And I write this from repeated experience: models that announce a hundred languages are usually worse in Arabic than models that announce twenty. Because a hundred languages means every language got whatever was left over, and twenty means there was a budget.

Second point: Arabic is the first name on the list.

That is English alphabetical order, not a priority ranking, and I claim nothing more. But the psychological and practical effect is real: anyone reading the model card sees Arabic first. And after thirty articles of hunting for Arabic in footnotes, that is a change you feel.

Third point, and this is a criticism I must direct at myself before anyone else.

I tried to extract a separate Arabic figure from this paper. I saw a table that appears to contain an Arabic column in a multilingual MMLU evaluation, but I could not verify it to my own satisfaction, so I excluded it. That is my rule across this entire series: a number I have not verified twice does not get published, however tempting it is.

And the uncomfortable result: even in a model that puts Arabic first on its list, I cannot give you a documented Arabic figure. Which returns us to the pattern we have seen in every article: announcing is easier than measuring, and Arabic measurement arrives late or not at all.


From the test notebook: evaluating a multilingual model on your own language

When no Arabic number exists, test it yourself. This is a method specific to multilingual models, different from what we used earlier.

Item one: the cross-language leakage test.

Multilingual models sometimes blend. Ask for pure Arabic and watch for:

What to look for An example of leakage
A Persian or Urdu word A shared lexeme used in its non-Arabic sense
Turkish or Hebrew structure An odd sentence order
Latin punctuation Latin commas and colons inside Arabic text
Mixed Eastern and Western digits “I have ٣ and 5 orders”

Ask for a three hundred word paragraph and count these. A clean multilingual model does not blend. One that does is telling you its Arabic sits beside other languages in the same space with no separation.

Item two: the Hebrew and Persian test.

This item is specific to this model, because its list includes Hebrew, Persian and Arabic together, and all three use closely related alphabets or the same writing direction.

Write an Arabic sentence containing a word that could be confused, then ask the model to identify its language:

The text The correct language The likely confusion
kitaab Arabic A near-identical word in Persian
dars Arabic Used in Urdu and Persian
nazar Arabic Has a different use in Persian
madrasa Arabic Used in several languages

A good model does not hesitate. A model that asks “did you mean Persian?” is confusing the alphabet with the language, which is a very common error in multilingual models.

Item three: the round-trip translation test.

Ask for a translation from Arabic into a third language on its list, then back into Arabic. For example Arabic to Turkish to Arabic.

The original: “lissa ma wasalni raddak, w iza ma beeji al-yom badtarr aʿtizer.”
(Your reply still has not reached me, and if it does not come today I will have to apologise.)

A good model returns something close in meaning. A weak one loses the conditional, or the negation, or the future apology, three precise grammatical elements. This test exposes a model’s depth of Arabic understanding better than any direct question.

Item four: the fair allocation test.

Ask the same knowledge question in three of its languages, one of them Arabic:

“Who was Ibn Khaldun and why is he considered important?”

Ask in Arabic, in English and in French. Compare length, depth and accuracy across the three. If the Arabic answer is shorter and shallower, the model knows about the Arab world in English better than it knows it in Arabic. That is a pattern I have seen often, and it is one of the strangest things in this field: knowledge about your culture, stored in somebody else’s language.

Item five: the language share test.

The depth-versus-breadth idea is practically testable. Compare two models of similar size, one declaring twenty languages and one declaring a hundred.

The task What it measures
A 300 word paragraph in Arabic Production quality
Morphological analysis of five words Grammatical depth
Converting a passage into Levantine Dialect
Five Arabic knowledge questions Knowledge

In my experience the difference shows in the last two rather than the first two. Surface production is something everyone manages. Depth requires a share.

Item six: the balance across the list test.

Ask the model the same question in each of its listed languages that you know, and compare. If its Arabic performance sits closer to the tail of the list than the head, Arabic’s presence on the list is administrative rather than engineering.

A practical method: ask for the same sentence translated into five of its languages, then translate all five back into Arabic. The language whose round trip comes back most accurately is the one it genuinely commands.


Depth against breadth: what three cases taught me

The case The choice The Arabic result
Jais Arabic first, bilingual The best open Arabic model of its time
AceGPT Arabic on top of a global base Won on style, stumbled on knowledge
Aya 23 Twenty three languages, Arabic among them Arabic as a principal language, not a marginal one

Three different paths and three different results. And none of them is wrong.

The conclusion I reach: there is no single correct path for getting Arabic into models. There is one correct question that should be asked at the start of every project: what share of this model is Arabic, in numbers? And anyone who cannot answer has not yet made a decision.


The practical verdict

Prefer the model that declares twenty languages over the one that declares a hundred. A short list is a promise. A long list is coverage.

And always ask for a separate figure for your language. Your language being on a list does not mean a measurement exists for it. Those are two different things and they are constantly conflated.

And if you cannot find the figure, assume neither the worst nor the best. Test. Half a day of testing saves you months of argument.


Next in the series

Two weeks later Alibaba released Qwen2, and its language list contains a group explicitly labelled “Middle East”, holding Arabic, Persian, Hebrew and Turkish. The next article is about a Chinese model that placed Arabic in a named regional group, and about the one published figure for its Arabic performance.


Sources

  • Cohere, C4AI Launches Aya 23, May 23, 2024: cohere.com/blog/aya23
  • Aryabumi et al., Aya 23: Open Weight Releases to Further Multilingual Progress, arXiv:2405.15032, May 23, 2024
  • Aya 23 technical report: cohere.com/research/aya