The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Number of languages | 23 | The model’s scope | Aya 23 technical report |
| Arabic’s position on the list | First | Alphabetical in English, but named explicitly | Same report |
| The languages | Arabic, Chinese, Czech, Dutch, English, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Korean, Persian, Polish, Portuguese, Romanian, Russian, Spanish, Turkish, Ukrainian, Vietnamese | Who is in and who is out | Same report |
| Its predecessor Aya 101 | 101 languages | The comparison point | Same report |
| Sizes | 8 and 35 billion parameters | Accessibility | Same report |
| The declared thesis | “An experiment in depth vs breadth, exploring the impact of allocating more capacity to fewer languages” | The core of the paper, quoted verbatim | Paper abstract |
| Release date | May 23, 2024 | Confirmed on arXiv and the Cohere blog | arXiv:2405.15032 |
| A separate Arabic figure I could confirm | None | I tried three times and could not establish one well enough to publish | My own methodological note |
That last row matters, and I will not hide it: a model that puts Arabic first on its list, and I could not confirm a separate Arabic number for it that I trust.
Depth against breadth
This chart illustrates the idea the paper argues, not a published measurement. And the idea is simple and strong: one language out of twenty three gets more than one language out of a hundred and one.
Model card
| Item | Detail |
|---|---|
| Model | Aya 23, at 8 and 35 billion parameters |
| Lab | Cohere For AI and Cohere |
| Release date | May 23, 2024 |
| Paper | arXiv:2405.15032 |
| Licence | Open weights |
| Arabic | Explicitly named among twenty three languages |
| What was published about Arabic | Its presence on the list, which is confirmed |
| What I could not confirm | Any separate Arabic performance figure |
The ARAB-LENS reading: why a short list matters to you more than a long one
This article is about an idea more than a number, and the idea earns its place.
This model’s predecessor, Aya 101, covered a hundred and one languages. At first glance that looks like an enormous achievement. A hundred languages! Coverage of the whole world!
Then the same team comes back in under a year and says: we tried the opposite. Twenty three languages only, and more capacity for each. The abstract calls it, in its own words, “an experiment in depth vs breadth”.
Why does that matter to an Arabic speaker specifically?
Because in long lists Arabic is treated as a minority language, and in short lists it is treated as a principal one. The difference is not in the naming but in the allocation: more data, sharper evaluation, native reviewers, and a heavier weight in training decisions.
And I write this from repeated experience: models that announce a hundred languages are usually worse in Arabic than models that announce twenty. Because a hundred languages means every language got whatever was left over, and twenty means there was a budget.
Second point: Arabic is the first name on the list.
That is English alphabetical order, not a priority ranking, and I claim nothing more. But the psychological and practical effect is real: anyone reading the model card sees Arabic first. And after thirty articles of hunting for Arabic in footnotes, that is a change you feel.
Third point, and this is a criticism I must direct at myself before anyone else.
I tried to extract a separate Arabic figure from this paper. I saw a table that appears to contain an Arabic column in a multilingual MMLU evaluation, but I could not verify it to my own satisfaction, so I excluded it. That is my rule across this entire series: a number I have not verified twice does not get published, however tempting it is.
And the uncomfortable result: even in a model that puts Arabic first on its list, I cannot give you a documented Arabic figure. Which returns us to the pattern we have seen in every article: announcing is easier than measuring, and Arabic measurement arrives late or not at all.
From the test notebook: evaluating a multilingual model on your own language
When no Arabic number exists, test it yourself. This is a method specific to multilingual models, different from what we used earlier.
Item one: the cross-language leakage test.
Multilingual models sometimes blend. Ask for pure Arabic and watch for:
| What to look for | An example of leakage |
|---|---|
| A Persian or Urdu word | A shared lexeme used in its non-Arabic sense |
| Turkish or Hebrew structure | An odd sentence order |
| Latin punctuation | Latin commas and colons inside Arabic text |
| Mixed Eastern and Western digits | “I have ٣ and 5 orders” |
Ask for a three hundred word paragraph and count these. A clean multilingual model does not blend. One that does is telling you its Arabic sits beside other languages in the same space with no separation.
Item two: the Hebrew and Persian test.
This item is specific to this model, because its list includes Hebrew, Persian and Arabic together, and all three use closely related alphabets or the same writing direction.
Write an Arabic sentence containing a word that could be confused, then ask the model to identify its language:
| The text | The correct language | The likely confusion |
|---|---|---|
| kitaab | Arabic | A near-identical word in Persian |
| dars | Arabic | Used in Urdu and Persian |
| nazar | Arabic | Has a different use in Persian |
| madrasa | Arabic | Used in several languages |
A good model does not hesitate. A model that asks “did you mean Persian?” is confusing the alphabet with the language, which is a very common error in multilingual models.
Item three: the round-trip translation test.
Ask for a translation from Arabic into a third language on its list, then back into Arabic. For example Arabic to Turkish to Arabic.
The original: “lissa ma wasalni raddak, w iza ma beeji al-yom badtarr aʿtizer.”
(Your reply still has not reached me, and if it does not come today I will have to apologise.)
A good model returns something close in meaning. A weak one loses the conditional, or the negation, or the future apology, three precise grammatical elements. This test exposes a model’s depth of Arabic understanding better than any direct question.
Item four: the fair allocation test.
Ask the same knowledge question in three of its languages, one of them Arabic:
“Who was Ibn Khaldun and why is he considered important?”
Ask in Arabic, in English and in French. Compare length, depth and accuracy across the three. If the Arabic answer is shorter and shallower, the model knows about the Arab world in English better than it knows it in Arabic. That is a pattern I have seen often, and it is one of the strangest things in this field: knowledge about your culture, stored in somebody else’s language.
Item five: the language share test.
The depth-versus-breadth idea is practically testable. Compare two models of similar size, one declaring twenty languages and one declaring a hundred.
| The task | What it measures |
|---|---|
| A 300 word paragraph in Arabic | Production quality |
| Morphological analysis of five words | Grammatical depth |
| Converting a passage into Levantine | Dialect |
| Five Arabic knowledge questions | Knowledge |
In my experience the difference shows in the last two rather than the first two. Surface production is something everyone manages. Depth requires a share.
Item six: the balance across the list test.
Ask the model the same question in each of its listed languages that you know, and compare. If its Arabic performance sits closer to the tail of the list than the head, Arabic’s presence on the list is administrative rather than engineering.
A practical method: ask for the same sentence translated into five of its languages, then translate all five back into Arabic. The language whose round trip comes back most accurately is the one it genuinely commands.
Depth against breadth: what three cases taught me
| The case | The choice | The Arabic result |
|---|---|---|
| Jais | Arabic first, bilingual | The best open Arabic model of its time |
| AceGPT | Arabic on top of a global base | Won on style, stumbled on knowledge |
| Aya 23 | Twenty three languages, Arabic among them | Arabic as a principal language, not a marginal one |
Three different paths and three different results. And none of them is wrong.
The conclusion I reach: there is no single correct path for getting Arabic into models. There is one correct question that should be asked at the start of every project: what share of this model is Arabic, in numbers? And anyone who cannot answer has not yet made a decision.
The practical verdict
Prefer the model that declares twenty languages over the one that declares a hundred. A short list is a promise. A long list is coverage.
And always ask for a separate figure for your language. Your language being on a list does not mean a measurement exists for it. Those are two different things and they are constantly conflated.
And if you cannot find the figure, assume neither the worst nor the best. Test. Half a day of testing saves you months of argument.
Next in the series
Two weeks later Alibaba released Qwen2, and its language list contains a group explicitly labelled “Middle East”, holding Arabic, Persian, Hebrew and Turkish. The next article is about a Chinese model that placed Arabic in a named regional group, and about the one published figure for its Arabic performance.
Sources
- Cohere, C4AI Launches Aya 23, May 23, 2024: cohere.com/blog/aya23
- Aryabumi et al., Aya 23: Open Weight Releases to Further Multilingual Progress, arXiv:2405.15032, May 23, 2024
- Aya 23 technical report: cohere.com/research/aya