The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Arabic among the optimised languages | Yes, by name | An explicit commitment from a Western lab | Model card, Cohere |
| Number of optimised languages | 10 | English, French, Spanish, Italian, German, Brazilian Portuguese, Japanese, Korean, Simplified Chinese, Arabic | Same card |
| Languages present in pretraining but not optimised | 13 | Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, Persian | Same card |
| Arabic result published by Cohere | None | No figure from the company itself | Card and documentation |
| Independent Arabic result, acc_norm | 48.91% | Aggregate score on the Open Arabic LLM Leaderboard | OALL results dataset, run of May 19, 2024 |
| Independent Arabic result, acc | 72.93% | The same evaluation in a different unit | Same source |
| Standard error, acc_norm | 0.0375 | Width of the confidence margin | Same source |
| Standard error, acc | 0.0114 | Much narrower | Same source |
| Run date | May 19, 2024 | Six weeks after release | Same source |
Note rows five and six: 48.91% and 72.93% for the same model on the same run. That is not a contradiction, it is a lesson in measurement that I will explain.
The language window
The difference between the first row and the second is the difference between “we take responsibility for it” and “it may work”.
Model card
| Item | Detail |
|---|---|
| Model | Command R+ |
| Lab | Cohere |
| Release date | April 4, 2024 |
| Orientation | Enterprise models, with a focus on retrieval and tool use |
| Arabic | On the list of optimised languages |
| Important warning | The August 2024 refresh is a different model, do not conflate them |
The ARAB-LENS reading: why this list matters more than any score
I will say it plainly: this is the first time I saw a major Western company write “Arabic” inside a commitment list rather than inside a footnote.
Review with me what we have seen so far in this series. Llama 2: Arabic under the table’s threshold. Falcon 180B: not mentioned. Mixtral: five languages and Arabic is not among them. Claude 3: named once, in preference data collection. And here, suddenly: Arabic written by name in a list of ten languages the model is “optimised to perform well in”.
Why do I consider that more important than a high benchmark score? Three practical reasons.
First, it is a contractual promise rather than an observation. When a company puts a language on an “optimised for” list, it attaches its name to it. You can open a support ticket and say “its Arabic performance is below standard” and point to their own document. With a model outside the list, you have nothing to point to.
Second, it means a resource decision. A language is not added to that list for free: evaluation data, native reviewers, alignment work. Arabic being on the list means somebody inside that company argued for it in a planning meeting.
Third, and most important for the region, it is a precedent. After Command R+ it became normal to ask other companies: where is Arabic on your list? That question was not askable before somebody answered it in the affirmative.
Now the other side, which is necessary.
Cohere itself published not one Arabic figure. The card states outright that its published metrics do not cover multilingual performance. So the commitment is declared and the measurement is missing. Which brings us back to the same pattern: companies announce, and academics and the community measure.
The only available figure came from the Open Arabic LLM Leaderboard, in a run dated May 19, 2024. And here is the measurement lesson I promised you.
A lesson in measurement: why 48.91 and 72.93 are both correct
This deserves its own section because I have watched many people trip over it.
In multiple-choice evaluation of language models there are two common metrics:
| Metric | How it is computed | When it is used |
|---|---|---|
| acc | Did the model pick the correct answer? | When the options are of similar length |
| acc_norm | The same, after normalising each option’s probability by its length | When option lengths differ |
Why normalise? Because a model naturally favours shorter options, since a short string has a mathematically higher probability. So if the correct answer is long, the model may reject it for a purely arithmetic reason unrelated to knowledge. Normalisation corrects that.
The result is that acc_norm is usually harsher and more honest, while acc is more forgiving. The difference between 48.91 and 72.93 here is enormous, and it tells us that option lengths in this benchmark are badly unbalanced to begin with.
Take this as a rule: when somebody gives you a number from an Arabic leaderboard, ask which metric. Anyone who does not know does not know what they are quoting. And note the standard error too: 0.0375 against 0.0114. The first metric’s confidence margin is three times wider, meaning it is less stable.
From the test notebook: verifying an “optimised for Arabic” claim
When a vendor promises you Arabic, you have a right to check. Four items that test the promise rather than general ability.
Item one: does the optimisation cover dialect or only formal Arabic?
Ask the same question in Modern Standard Arabic and in three dialects, and compare the quality of the replies:
| Register | The text |
|---|---|
| Formal | “arghab fee ilghaaʾ talabi wa istirdaad al-mablagh.” |
| Levantine | “biddi alghi talabiyyti w arajjeʿ masaariyyi.” |
| Egyptian | “ʿaayez alghi al-order w arrajjaʿ feloosi.” |
| Gulf | “abi alghi talabi w astaridd feloosi.” |
Companies that “optimise for Arabic” usually optimise for formal Arabic, because that is where the evaluation data exists. If the reply is excellent on the first and weak on the other three, the promise is true but narrower than you understood.
Item two: the direction and punctuation test.
A simple technical item that exposes models nobody polished:
| Case | Correct output | Common error |
|---|---|---|
| A number inside an Arabic sentence | “ʿandi 3 talabaat” | Digit order reverses |
| English text inside Arabic | “iftah malaff report.pdf” | The dot and extension jump position |
| Parentheses around Arabic | “(al-talab mulgha)” | Opening and closing brackets flip |
| Arabic comma | The Arabic comma | Uses the Latin one |
| Percent sign | After the number | Placed before it |
These are small errors, but they show up instantly in a user interface and make the product look translated rather than built.
Item three: the answer length test.
Ask for a fifty word answer in Arabic, then in English. Models not tuned for Arabic produce Arabic far longer than requested, because they count tokens rather than words and the letter tax confuses them. If you asked for fifty and got a hundred and twenty, the model is not measuring its own length in Arabic.
Item four: consistency across two sessions.
Ask the same Arabic question today and again in a week. A genuinely optimised model gives an answer of comparable quality. A model that merely encounters Arabic fluctuates. This item is slow but it is the most honest.
Item five: the Arabic retrieval test.
Command R+ is built for retrieval, so test it in its speciality:
| The task | What it measures |
|---|---|
| Give it three Arabic documents and ask a question requiring two | Synthesis across sources |
| Place contradictory information in two documents | Contradiction detection |
| Ask about information absent from the documents | Admitting absence |
| Ask for a verbatim quotation from a document | Fidelity |
The third item is the most important and the most failed: a model that invents an answer from its own knowledge rather than saying “not in the documents” ruins any Arabic retrieval system.
Item six: the tool and function test.
Enterprise models call tools. Test that in Arabic:
“Book me an appointment with Doctor Samir next Thursday at three in the afternoon.”
Watch: did it extract the name, the date and the time correctly? “Next Thursday” is a relative date, and “three in the afternoon” needs converting to 15:00. Models that handle this well in English frequently err in Arabic, especially with relative and dialectal dates.
Optimised-language lists: five companies compared
I gathered what the companies published up to this date into one table:
| Company and model | Declared languages | Arabic |
|---|---|---|
| Mistral, Mixtral 8x7B | 5 | No |
| Meta, Llama 2 | No list, a general warning | No |
| TII, Falcon 180B | 4 full and 7 limited | No |
| Anthropic, Claude 3 | No optimised-language list | No |
| Cohere, Command R+ | 10 | Yes |
One of five. That is the size of what had been achieved by April 2024.
Why do I insist on tracking this rather than tracking scores? Because a score changes with every release, while a list reflects an institutional decision that persists. A company that put your language on its list once will put it there again. A company that has not needs a reason to start. And that reason is the market, not linguistic justice.
The practical verdict
The optimised-languages list is the first thing I read in any model card, before any number. That is the single most useful habit I picked up from writing this series.
And never quote a figure from an Arabic leaderboard without naming the metric and the date. “Scored 48.91 on acc_norm on the OALL leaderboard in May 2024” is a defensible sentence. “Scored 48.91 in Arabic” is an incomplete one.
Watch the version names. Command R+ of April 2024 and Command R+ of August 2024 are different models. That confusion happens constantly and it ruins any comparison over time.
Next in the series
Forty days later, OpenAI changed its tokenizer for the first time since ChatGPT. And the result for Arabic was a specific, published number: from 53 tokens down to 26 for the same sentence. The next article is about the day the letter tax was cut in half.
Sources
- Cohere, Introducing Command R+, April 4, 2024: cohere.com/blog/command-r-plus-microsoft-azure
- Model card c4ai-command-r-plus on Hugging Face
- Open Arabic LLM Leaderboard results dataset, run file of May 19, 2024: huggingface.co/datasets/OALL/results
- Hugging Face and TII, Introducing the Open Arabic LLM Leaderboard, May 2024: huggingface.co/blog/leaderboard-arabic