The numbers first
| Metric | Result | What it measures | Source |
|---|---|---|---|
| Model size | 180 billion parameters | The largest open model of its moment | Technology Innovation Institute press release, September 6, 2023 |
| Training data | 3.5 trillion tokens | Also the largest | Same release |
| Languages stated on the model card | English, German, Spanish, French, with limited ability in seven more | What the lab itself declares | Model card on Hugging Face |
| Arabic mentioned on the model card | Not present | Whether Arabic is officially supported | Same card |
| Arabic mentioned in the press release | Not present | Whether it was mentioned at all | TII release |
| RefinedWeb English share | 75% | The largest block | Model card |
| RefinedWeb Europe share | 7% | The other languages, and they are European | Same card |
| Published Arabic score for Falcon 180B | None exists | I searched five sources and found nothing | ArabicMMLU, AlGhafa, the OALL leaderboard, a 2025 survey |
| Falcon 40B on ArabicMMLU, for comparison | 34.8% | Against 29.0% for random guessing | Koto et al., February 2024 |
| Falcon-instruct 40B on ArabicMMLU | 33.0% | Lower than the base version | Same source |
The important row is the one with no number in it: there is no published Arabic result for the largest open model an Arab country ever built.
The data mix, as the institute published it
“RefinedWeb Europe” is specifically the European-language subset. Arabic is not a line in this distribution, and no share is declared for it.
Model card
| Item | Detail |
|---|---|
| Model | Falcon 180B |
| Institution | Technology Innovation Institute, Abu Dhabi |
| Release date | September 6, 2023 |
| Scale | 180 billion parameters, 3.5 trillion tokens |
| Standing at the time | Number one in the world among open models |
| Data paper | RefinedWeb, arXiv:2306.01116, June 2023 |
| What the RefinedWeb paper says | “For this paper, we focus on English” |
| The publicly released dataset | English only |
The ARAB-LENS reading: why the absence is the story
I will say the sentence I think sums up a decade of technology policy in the region: an Arab country built the largest open model in the world and did not put Arabic in it.
Before I am misread, three clarifications.
First, this was not stupidity. It was a clear strategic decision. The institute wanted to top the global leaderboards, and those leaderboards are all in English. Had they spent twenty percent of their compute on Arabic they would have dropped places in the global ranking and nobody would have written about them. The commercial arithmetic was sound.
Second, the UAE itself released Jais one week earlier. So the region had two models in seven days: one mid-sized and Arabic, one enormous and Arabic-free. That is a conscious separation of two tracks, not neglect.
Third, and here is my real criticism: the problem is not the decision, it is the story told about it. Falcon was presented across Arab media as an Arab achievement, and the image travelled as proof that the region had reached the front rank. Nobody told the Arab reader that this model does not speak their language. And when I went looking for any figure telling me how much Arabic Falcon 180B knows, I found none. Not from the institute and not from any third party.
The only number available is for its smaller sibling: Falcon 40B scored 34.8% on ArabicMMLU. Random guessing on that exam is 29.0%. Which means the most famous Arab-built model in the world beat a dice roll by fewer than six points.
And there is a harsher detail. The instruction-tuned version, Falcon-instruct 40B, scored 33.0%, which is lower than the base version. Instruction tuning, a stage that is supposed to improve a model, made it worse in Arabic. That happens when the tuning data is entirely English.
An important methodological note: I have seen a figure circulating that Arabic made up about 1.2% of RefinedWeb. That figure belongs to a multilingual pipeline derived from CommonCrawl, not to Falcon 180B’s training mix, and it must not be presented as the Arabic share of the model. I mention it here because I want you to know I saw it and deliberately left it out.
From the test notebook: measuring a model that does not claim Arabic
When a model is not built for Arabic, the test questions change. Do not ask “is it good?”. Ask “where exactly does it break?”.
Item one: the collapse ladder.
Ask for the same task at five linguistic levels and record where the failure starts:
| Level | The task | What it measures |
|---|---|---|
| 1 | Translate “good morning” into Arabic | Basic vocabulary |
| 2 | Write a sentence about the weather in Arabic | Simple structure |
| 3 | Write a paragraph about the history of Damascus | Arabic knowledge |
| 4 | Explain the difference between the absolute object and the object of purpose | Advanced grammar |
| 5 | Convert this paragraph from Modern Standard Arabic into Levantine | Dialect |
Models not built for Arabic succeed at 1 and 2, stumble at 3, invent at 4, and fail completely at 5. Recording the level number is far more useful than recording a pass or a fail.
Item two: the geographic hallucination test.
Ask five questions about specific Arab places:
| The question | Why it is a trap |
|---|---|
| “How far is Irbid from Amman?” | Cities with no global fame |
| “What is the capital of Dhi Qar governorate?” | Iraqi administrative divisions |
| “Where is Souq Waqif?” | A Qatari landmark |
| “What is the climate difference between Jeddah and Medina?” | Saudi geography |
| “Which river runs through Deir ez-Zor?” | Syrian geography |
A model weak in Arabic will answer the first and third and confidently invent the rest. The confidence is the problem, not the ignorance.
Item three: the professional term test.
Every Arab market has its own vocabulary. Examples from different fields:
| The term | The field | The meaning |
|---|---|---|
| kafaala | Gulf employment | The sponsorship system for a migrant worker |
| taabu | Levantine real estate | The title deed |
| baraaʾat dhimma | Administration | A certificate of no outstanding obligations |
| hawaala | Finance | A trust-based money transfer |
| mukhaalasa | Employment | The final settlement on leaving a job |
Ask about each. A model that translates them literally into English and then explains the English meaning did not learn them, it inferred them. And this is the most dangerous error class in any commercial application, because it looks correct.
Item four: the recent Arab history test.
General historical knowledge is available in English, but close detail is not:
| The question | What it measures |
|---|---|
| “When did Jordan gain independence and who was its first king?” | Basic national history |
| “What was Syria’s currency before the lira?” | A local economic detail |
| “Who built Ajloun Castle and why?” | Regional history |
| “What was the Taif Agreement?” | Modern Arab politics |
A model that answers the first and fourth and invents the second and third knows what was written about us in English, not what we know about ourselves.
Item five: the weights and measures test.
A simple and very revealing item:
“I want to buy 3 kilos of meat and 2 ratls of cheese, at 12 dinars a kilo and 5 a ratl. What is the total?”
The ratl is a unit used in the Levant and Iraq whose value in kilos differs from country to country. A good model flags the ambiguity and asks. A weak one assumes a value and calculates confidently. And that species of confident wrongness is the most dangerous thing in any commercial application.
A strategic lesson from Falcon: how to read national model announcements
From this case I built a checklist I use with every “national model” or “Arabic model” announcement:
| The question | Why |
|---|---|
| Is Arabic on the declared language list? | The simplest and strongest check |
| Was the Arabic share of the data published? | If not, assume it is small |
| Is there a number on a native Arabic benchmark? | Not a translated one |
| Is the number for the announced model or its smaller sibling? | A very common conflation |
| Who funded and who measured? | The same body or an independent one? |
In the case of Falcon 180B the answers to all five are: no, no, no, for its smaller sibling, and no measurement exists at all.
I do not write this to attack a project, but because the region needs sharper readers of its own announcements. Applause without reading has cost us more than any technical shortfall.
The practical verdict
Do not use a model in an Arabic product when its own card does not mention Arabic. That is not a strict rule, it is a literal reading of what the lab says about its own model. A model card is a document, and when it does not name your language it is telling you something.
If you are forced to, test on the five levels above before any commitment, and set explicit boundaries: no dialect, no local knowledge, no professional terminology.
And if you are a journalist or a technical writer, always ask before you call a model “Arab”: Arab in origin or Arab in capability? The difference between the two is the difference between a car factory in your country and a car suited to your roads.
Next in the series
Two weeks later an academic team published a paper arguing that the problem is neither scale nor language but localisation. Their model, AceGPT, beat GPT-3.5 at Arabic instruction following at only thirteen billion parameters. Then lost to it on knowledge and on culture. The next article explains why.
Sources
- Technology Innovation Institute, Technology Innovation Institute Introduces World’s Most Powerful Open LLM: Falcon 180B, September 6, 2023: tii.ae
- Hugging Face, Spread Your Wings: Falcon 180B is here, September 6, 2023: huggingface.co/blog/falcon-180b
- Model card tiiuae/falcon-180B on Hugging Face
- Almazrouei et al., The RefinedWeb Dataset for Falcon LLM, arXiv:2306.01116, June 2023
- Koto et al., ArabicMMLU, Findings of ACL 2024: aclanthology.org/2024.findings-acl.334