The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Falcon 2 11B on Arabic ARC-C, 25-shot | 25.32 | A four-option test whose random baseline is 25 | Technical report, arXiv:2407.14885, Table 11 |
| On Arabic MMLU, 25-shot | 28.04 | General knowledge | Same table |
| On Arabic HellaSwag, zero-shot | 32.42 | Commonsense reasoning | Same table |
| On Arabic TruthfulQA, zero-shot | 49.66 | Resistance to falsehood | Same table |
| Languages declared on the model card | 11, none of them Arabic | English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech, Swedish | Model card tiiuae/falcon-11B |
| Arabic mentioned in the institute’s release | Not present | Whether it was named | TII press release |
| Arabic in the training data mixture | Not present | Table 3 of the report: every named language is European | Same report |
| English in the mixture | 61% to 69.1% | The largest block | Same table |
| “Multilingual” in the mixture | 15% to 16.6% | Ten European languages | Same table |
| Tables in which Arabic appears | One, in the whole report | The scale of attention | Same report |
The first row is the article: 25.32 on a test whose luck baseline is 25. Three tenths of a point above a dice roll.
The only Arabic number, next to the chance line
Three of four are touching the chance line. The fourth, TruthfulQA, is a different kind of measure and is not directly comparable, because a model that declines to answer often can score high on it.
Model card
| Item | Detail |
|---|---|
| Model | Falcon 2 11B |
| Institution | Technology Innovation Institute, Abu Dhabi |
| Release date | May 13, 2024 |
| Size | 11 billion parameters |
| Technical report | arXiv:2407.14885, July 2024 |
| Declared languages | Eleven, all European plus English |
| Arabic | Not on the card, not in the release, not in the training mixture |
| The only Arabic benchmark | The Open Multilingual LLM Leaderboard, which is machine-translated |
The ARAB-LENS reading: the second time, and from the same institution
In article twenty five I wrote about Falcon 180B and the absence of Arabic from it, and I said that was an understandable strategic decision: they wanted the top of the global leaderboards.
Now, eight months later, a new model, and the language list is still without Arabic. Eleven languages: English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech and Swedish.
Swedish is there. Czech is there. Romanian is there. Arabic is not.
I write that sentence without heat, as a documentary observation: Swedish has around ten million speakers, Arabic has hundreds of millions, and the institute that built the model sits in an Arab capital.
Then the training data table in the technical report confirms it: the “multilingual” component makes up about fifteen percent, and the languages named inside it are ten European ones. Arabic is not a line item in the training mixture at all.
So where did the number come from?
From a single table at the back of the report, headed as an evaluation on the Open Multilingual LLM Leaderboard. That leaderboard takes English benchmarks and machine-translates them into many languages, Arabic among them. Which means the only available figure is the result of a model never trained on Arabic, on an exam machine-translated into Arabic. Two layers of distance from the real language.
And yet the number is very useful, and that is what makes this article worth writing.
Why does 25.32 matter? Because ARC-C is a four-option multiple choice test. Pure random guessing yields twenty five percent on average. The model scored twenty five point three two. Which means its performance on the Arabic questions is statistically indistinguishable from chance.
That is not a criticism of the model. The model never claimed Arabic. It is a clean measurement of what happens when a language is not in the data. I consider this one of the most useful numbers in the whole series, because it gives you the floor: this is what an Arabic zero looks like when you measure it.
Keep it as a reference. The next time somebody tells you a model “supports Arabic”, ask for its score on a multiple choice benchmark. If it is near the chance line, you are looking at Falcon 2 wearing a different face.
From the test notebook: computing the chance line yourself
A short methodological section that will change how you read numbers forever.
The rule: chance line = 100 divided by the number of options.
| Test type | Options | Chance line |
|---|---|---|
| ARC-C, HellaSwag | 4 | 25.0% |
| MMLU | 4 | 25.0% |
| ArabicMMLU | Variable, published as 29.0% | 29.0% |
| True or false | 2 | 50.0% |
| Five-way choice | 5 | 20.0% |
The application: take any number you read and subtract the chance line. That is the real performance.
| Model and test | Published figure | Chance line | Real performance |
|---|---|---|---|
| Falcon 2 11B on Arabic ARC-C | 25.32 | 25.0 | 0.32 points |
| Falcon 2 11B on Arabic MMLU | 28.04 | 25.0 | 3.04 points |
| LLaMA2-7B on Arabic MMLU | 29.47 | 29.0 | 0.47 points |
| Falcon 40B on ArabicMMLU | 34.8 | 29.0 | 5.8 points |
| Jais-chat 30B on ArabicMMLU | 62.3 | 29.0 | 33.3 points |
| GPT-4 on ArabicMMLU | 72.5 | 29.0 | 43.5 points |
Look at the last column. Differences that look close in the first column become enormous in the last. The gap between Falcon 40B and Jais-chat 30B is not twenty seven points. It is thirty three point three against five point eight, which is more than five times.
This is the single most useful arithmetic habit I can recommend to anyone reading Arabic model numbers: always subtract the chance line before you compare.
A bonus item: the consistency check.
A fast way to verify a model is genuinely guessing: ask it the same question five times, shuffling the option order each time. A model that knows the answer picks the same content. A model that is guessing tends to pick the same position, or moves at random. That exposes guessing even when it happens to have scored respectably.
Item four: the calibrated confidence test.
A model that is guessing ought to know it is guessing. Ask for a confidence estimate with every answer:
“Answer the question, then give me your confidence in the answer from 0 to 100.”
Then gather twenty questions and compare. A well calibrated model: when it says 90% it is right about nine times in ten. A poorly calibrated one says 90% and is right half the time.
This item matters more in Arabic than elsewhere, because models not trained on it keep a confident tone even while guessing. And linguistic confidence deceives a user more than the error itself does.
Item five: the gradual collapse test.
Ask for the same task at five lengths: a sentence, a paragraph, three paragraphs, a page, two pages. Record where the text begins to fall apart.
| Length | What happens in a model weak in Arabic |
|---|---|
| Sentence | Sound |
| Paragraph | Usually sound |
| Three paragraphs | Repetition begins |
| A page | Clear repetition and topic drift |
| Two pages | Broken sentences or a switch to English |
The length at which it collapses is more useful than any benchmark score, because it tells you exactly what you can ask of it.
The Arabic floor: a permanent reference
I will close with a table worth saving, gathering every “near zero” figure we have seen in this series:
| The model | The test | The figure | Above chance by |
|---|---|---|---|
| Falcon 2 11B | Arabic ARC-C | 25.32 | 0.32 |
| LLaMA2-7B | Arabic MMLU | 29.47 | 0.47 |
| Falcon 2 11B | Arabic MMLU | 28.04 | 3.04 |
| LLaMA2-7B | Arabic EXAMs | 23.48 | below chance |
| Llama 2-Chat 7B | araSwag | 24.44 | below chance |
The last two rows deserve attention: performance below the chance line. That is not only ignorance, it is systematic bias: the model picks by a fixed wrong pattern, usually because it prefers a particular option by its position or its length.
Anyone who sees a figure below the chance line and reads it as “slightly weak” has not understood what they are reading.
The practical verdict
Read the model’s language list, then subtract the chance line from its numbers. Two steps are enough to filter ninety percent of claims.
And do not evaluate an Arabic model with machine-translated tests. The irony is that the only available figure for this model is a translated one, and I had to use it because nothing else exists, and I am telling you that outright rather than hiding it.
The strategic observation: the UAE produced Jais and Falcon in the same period. The first Arabic, the second global. That is a legitimate path, but it means Arabic will not reach the frontier models by accident. It arrives by decision, and the price is paid in global ranking.
Next in the series
One day after Falcon 2’s launch, the Open Arabic LLM Leaderboard was released by Hugging Face and the Technology Innovation Institute. And nine months later its own maintainers acknowledged three problems with the first version, one of them a scoring bug that affected the published ranking. The next article is about the leaderboard that corrected itself in public.
Sources
- Technology Innovation Institute, Falcon 2 series launch, May 13, 2024: tii.ae
- Hugging Face, Falcon 2 11B, May 2024: huggingface.co/blog/falcon2-11b
- Model card tiiuae/falcon-11B on Hugging Face
- Malartic, Roy Chowdhury et al., Falcon2-11B Technical Report, arXiv:2407.14885, July 2024, Tables 3 and 11