25.32 On a Test Whose Random Baseline Is 25: Falcon 2’s Only Arabic Number

Reading Time: 8 min
18
Ehab Saleh | techkahwa.net Model released: May 13, 2024 | Published: May 30, 2024

The numbers first

Metric Result What it measures Source
Falcon 2 11B on Arabic ARC-C, 25-shot 25.32 A four-option test whose random baseline is 25 Technical report, arXiv:2407.14885, Table 11
On Arabic MMLU, 25-shot 28.04 General knowledge Same table
On Arabic HellaSwag, zero-shot 32.42 Commonsense reasoning Same table
On Arabic TruthfulQA, zero-shot 49.66 Resistance to falsehood Same table
Languages declared on the model card 11, none of them Arabic English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech, Swedish Model card tiiuae/falcon-11B
Arabic mentioned in the institute’s release Not present Whether it was named TII press release
Arabic in the training data mixture Not present Table 3 of the report: every named language is European Same report
English in the mixture 61% to 69.1% The largest block Same table
“Multilingual” in the mixture 15% to 16.6% Ten European languages Same table
Tables in which Arabic appears One, in the whole report The scale of attention Same report

The first row is the article: 25.32 on a test whose luck baseline is 25. Three tenths of a point above a dice roll.


The only Arabic number, next to the chance line

Falcon 2 11B on translated Arabic benchmarks
ARC-C
25.32 chance line 25.00
MMLU
28.04 chance line 25.00
HellaSwag
32.42 chance line 25.00
TruthfulQA
49.66 a different kind of metric

Three of four are touching the chance line. The fourth, TruthfulQA, is a different kind of measure and is not directly comparable, because a model that declines to answer often can score high on it.


Model card

Item Detail
Model Falcon 2 11B
Institution Technology Innovation Institute, Abu Dhabi
Release date May 13, 2024
Size 11 billion parameters
Technical report arXiv:2407.14885, July 2024
Declared languages Eleven, all European plus English
Arabic Not on the card, not in the release, not in the training mixture
The only Arabic benchmark The Open Multilingual LLM Leaderboard, which is machine-translated

The ARAB-LENS reading: the second time, and from the same institution

In article twenty five I wrote about Falcon 180B and the absence of Arabic from it, and I said that was an understandable strategic decision: they wanted the top of the global leaderboards.

Now, eight months later, a new model, and the language list is still without Arabic. Eleven languages: English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech and Swedish.

Swedish is there. Czech is there. Romanian is there. Arabic is not.

I write that sentence without heat, as a documentary observation: Swedish has around ten million speakers, Arabic has hundreds of millions, and the institute that built the model sits in an Arab capital.

Then the training data table in the technical report confirms it: the “multilingual” component makes up about fifteen percent, and the languages named inside it are ten European ones. Arabic is not a line item in the training mixture at all.

So where did the number come from?

From a single table at the back of the report, headed as an evaluation on the Open Multilingual LLM Leaderboard. That leaderboard takes English benchmarks and machine-translates them into many languages, Arabic among them. Which means the only available figure is the result of a model never trained on Arabic, on an exam machine-translated into Arabic. Two layers of distance from the real language.

And yet the number is very useful, and that is what makes this article worth writing.

Why does 25.32 matter? Because ARC-C is a four-option multiple choice test. Pure random guessing yields twenty five percent on average. The model scored twenty five point three two. Which means its performance on the Arabic questions is statistically indistinguishable from chance.

That is not a criticism of the model. The model never claimed Arabic. It is a clean measurement of what happens when a language is not in the data. I consider this one of the most useful numbers in the whole series, because it gives you the floor: this is what an Arabic zero looks like when you measure it.

Keep it as a reference. The next time somebody tells you a model “supports Arabic”, ask for its score on a multiple choice benchmark. If it is near the chance line, you are looking at Falcon 2 wearing a different face.


From the test notebook: computing the chance line yourself

A short methodological section that will change how you read numbers forever.

The rule: chance line = 100 divided by the number of options.

Test type Options Chance line
ARC-C, HellaSwag 4 25.0%
MMLU 4 25.0%
ArabicMMLU Variable, published as 29.0% 29.0%
True or false 2 50.0%
Five-way choice 5 20.0%

The application: take any number you read and subtract the chance line. That is the real performance.

Model and test Published figure Chance line Real performance
Falcon 2 11B on Arabic ARC-C 25.32 25.0 0.32 points
Falcon 2 11B on Arabic MMLU 28.04 25.0 3.04 points
LLaMA2-7B on Arabic MMLU 29.47 29.0 0.47 points
Falcon 40B on ArabicMMLU 34.8 29.0 5.8 points
Jais-chat 30B on ArabicMMLU 62.3 29.0 33.3 points
GPT-4 on ArabicMMLU 72.5 29.0 43.5 points

Look at the last column. Differences that look close in the first column become enormous in the last. The gap between Falcon 40B and Jais-chat 30B is not twenty seven points. It is thirty three point three against five point eight, which is more than five times.

This is the single most useful arithmetic habit I can recommend to anyone reading Arabic model numbers: always subtract the chance line before you compare.

A bonus item: the consistency check.

A fast way to verify a model is genuinely guessing: ask it the same question five times, shuffling the option order each time. A model that knows the answer picks the same content. A model that is guessing tends to pick the same position, or moves at random. That exposes guessing even when it happens to have scored respectably.

Item four: the calibrated confidence test.

A model that is guessing ought to know it is guessing. Ask for a confidence estimate with every answer:

“Answer the question, then give me your confidence in the answer from 0 to 100.”

Then gather twenty questions and compare. A well calibrated model: when it says 90% it is right about nine times in ten. A poorly calibrated one says 90% and is right half the time.

This item matters more in Arabic than elsewhere, because models not trained on it keep a confident tone even while guessing. And linguistic confidence deceives a user more than the error itself does.

Item five: the gradual collapse test.

Ask for the same task at five lengths: a sentence, a paragraph, three paragraphs, a page, two pages. Record where the text begins to fall apart.

Length What happens in a model weak in Arabic
Sentence Sound
Paragraph Usually sound
Three paragraphs Repetition begins
A page Clear repetition and topic drift
Two pages Broken sentences or a switch to English

The length at which it collapses is more useful than any benchmark score, because it tells you exactly what you can ask of it.


The Arabic floor: a permanent reference

I will close with a table worth saving, gathering every “near zero” figure we have seen in this series:

The model The test The figure Above chance by
Falcon 2 11B Arabic ARC-C 25.32 0.32
LLaMA2-7B Arabic MMLU 29.47 0.47
Falcon 2 11B Arabic MMLU 28.04 3.04
LLaMA2-7B Arabic EXAMs 23.48 below chance
Llama 2-Chat 7B araSwag 24.44 below chance

The last two rows deserve attention: performance below the chance line. That is not only ignorance, it is systematic bias: the model picks by a fixed wrong pattern, usually because it prefers a particular option by its position or its length.

Anyone who sees a figure below the chance line and reads it as “slightly weak” has not understood what they are reading.


The practical verdict

Read the model’s language list, then subtract the chance line from its numbers. Two steps are enough to filter ninety percent of claims.

And do not evaluate an Arabic model with machine-translated tests. The irony is that the only available figure for this model is a translated one, and I had to use it because nothing else exists, and I am telling you that outright rather than hiding it.

The strategic observation: the UAE produced Jais and Falcon in the same period. The first Arabic, the second global. That is a legitimate path, but it means Arabic will not reach the frontier models by accident. It arrives by decision, and the price is paid in global ranking.


Next in the series

One day after Falcon 2’s launch, the Open Arabic LLM Leaderboard was released by Hugging Face and the Technology Innovation Institute. And nine months later its own maintainers acknowledged three problems with the first version, one of them a scoring bug that affected the published ranking. The next article is about the leaderboard that corrected itself in public.


Sources

  • Technology Innovation Institute, Falcon 2 series launch, May 13, 2024: tii.ae
  • Hugging Face, Falcon 2 11B, May 2024: huggingface.co/blog/falcon2-11b
  • Model card tiiuae/falcon-11B on Hugging Face
  • Malartic, Roy Chowdhury et al., Falcon2-11B Technical Report, arXiv:2407.14885, July 2024, Tables 3 and 11