Eight Supported Languages, and Arabic Is Not One: Llama 3.1 at 405 Billion

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Model released: July 23, 2024 | Published: August 6, 2024

The numbers first

Metric Result What it measures Source
Officially supported languages Eight: English, German, French, Italian, Portuguese, Hindi, Spanish, Thai Meta’s declared commitment Llama 3.1 model card
Arabic among them No Whether your language is on the list Same card
Multilingual share of the data Roughly 8% of the mix The share of every non-English language combined The Llama 3 Herd of Models paper
The full mix 50% general knowledge, 25% mathematical and reasoning, 17% code, 8% multilingual Where the capacity went Same paper, quoted verbatim
Total training tokens About 15 trillion Scale Model card
Languages documents were classified into 176 A fastText language identification model Same paper
Multilingual MMLU for the 405B model Portuguese 84.95, Spanish 85.08, Italian 85.04, German 84.36, French 84.66, Hindi 80.31, Thai 78.21 Seven languages, no Arabic Model card
Published Arabic figure for Llama 3.1 405B None exists Not from Meta and not from any third party Card, paper, Global MMLU, AraLingBench
Knowledge cutoff December 2023 The edge of what it knows Model card

The second to last row is the news: the largest open model in the world, and no published Arabic figure for it from anyone.


Where the fifteen trillion went

Llama 3 training data mix
General knowledge
50%
Maths and reasoning
25%
Code
17%
Multilingual
8%

Eight percent for every non-English language on earth. That division, rather than the language itself, is what determines how well these models will understand you.


Model card

Item Detail
Model Llama 3.1, at 8, 70 and 405 billion parameters
Lab Meta
Release date July 23, 2024
Paper The Llama 3 Herd of Models, arXiv:2407.21783
Training data About 15 trillion tokens of publicly available data
Supported languages Eight
Meta’s position on the rest “Llama may be able to output text in other languages than those that meet performance thresholds for safety and helpfulness”
Licence Explicitly permits fine-tuning for languages beyond the eight

The ARAB-LENS reading: what changed since Llama 2, and what did not

In article twenty two I wrote about Llama 2: Arabic under the data table’s threshold, meaning less than five parts in a hundred thousand. Now, a year later, where do we stand?

What changed, and it is a real change:

The mix now contains an explicit multilingual line at eight percent. In a model trained on fifteen trillion tokens, eight percent means about one trillion two hundred billion non-English tokens. That is an enormous figure by any measure, and many times more than anything available in the previous generation.

Meta also now classifies its documents into a hundred and seventy six languages using a language identification model. Which means Arabic at least gets seen in the pipeline now, rather than falling into an “unknown” bucket.

What did not change:

The official list is eight languages, and Arabic is not among them. And the multilingual MMLU table in the model card presents seven languages, Arabic not among those either.

Note the composition: Portuguese, Spanish, Italian, German, French, Hindi and Thai. Seven languages, five of them European. Thai is there and Arabic is not. I do not write that as an objection but as documentation of a visible ordering of priorities.

Now the arithmetic I want you to see.

One trillion two hundred billion tokens distributed across the non-English languages. If it were divided equally across a hundred and seventy six languages, each would receive about six billion eight hundred million tokens. That is far more than anything we saw in Llama 2.

But the division is not equal, and that is the heart of the matter. The official list is eight languages, and those are the ones that meet the “performance thresholds”. So the largest share went to the seven non-English languages on the list, and the remainder was split among everyone else. I put this arithmetic in front of you as an illustrative hypothesis and not as a published figure, because Meta did not publish the distribution.

The final point, and the most practically important: there is no Arabic figure for this model from any source. I searched the model card, the paper, the Global MMLU paper which evaluated the 8B and 70B versions but not the 405B, and AraLingBench which covered models only up to seventy billion. Nothing.

That deserves a moment’s reflection. The largest open model produced up to that date, available to everyone under an open licence, and nobody published how much Arabic it knows. Not because anyone prevented the measurement, but because the model’s size makes running it expensive, and Arab researchers measure what they can afford to run.

That, in my view, is one of the most dangerous gaps in the field: measurement itself has become a privilege that requires a compute budget.


From the test notebook: what to do when you cannot measure the model yourself

A common practical situation: the model is enormous and you do not have the hardware. This is my approach.

Item one: measure the smaller sibling and read the direction.

If a model family comes in three sizes, measure the small and the medium. The result does not give you the large model’s number, but it gives you the direction: does Arabic improve with scale in this family, or stay flat?

The pattern you see What it means
Large improvement from 8B to 70B Arabic is in the data and scale is extracting it
Slight improvement Arabic is nearly absent, and scale cannot create it
Degradation The Arabic data is contaminated or translated

Item two: test through a hosting provider.

Large models are available through hosting platforms at reasonable prices. Fifty questions from your own exam cost less than a cup of coffee. You do not need to run the model to measure it.

Item three: treat fine-tuning as an option.

The Llama 3.1 licence explicitly permits fine-tuning for languages beyond the eight. That opens a practical path: a model with a strong base and weak Arabic, which you tune on your own data. And that is exactly what many Arab teams did.

But note the limit: fine-tuning improves style, format and domain, and does not plant new knowledge at density. That is a lesson we saw with numbers in the AceGPT article.

Item four: use the eight languages as a reference scale.

A trick I use constantly. Ask the same question in Arabic and in Spanish, which is on the supported list. Then compare.

“Briefly explain three differences between a lease agreement and a sale agreement.”

The gap you see between the two answers is the cost of not being supported, measured directly. It is a fast and practical instrument that needs no benchmark and no leaderboard. I use it every time I evaluate a model whose list does not include my language.

Item five: the difference between sizes in the same family.

A methodological item I recommend to anyone choosing within a model family:

The comparison What it reveals
8B against 70B in English The gain from scale in general
8B against 70B in Arabic The gain from scale in your language
The ratio between the two gains Whether your language benefits from scale as others do

If the gain from scale is ten points in English and two in Arabic, scale does not solve your problem. Which means increasing your budget will not improve your Arabic product, and that you should be spending on data rather than on compute.

That, in my view, is the single most valuable piece of information an Arab team can obtain before committing a budget.

Item six: the Arabic safety test.

Meta ties supported languages to “performance thresholds for safety and helpfulness”. Which implies that Arabic did not meet the safety threshold. Test that:

The request What you watch
A general medical question in Arabic Does it add an appropriate caution?
A local legal question Does it flag that systems differ?
Culturally sensitive content Careful handling or arbitrary handling?
An attempted bypass in Arabic Are the guardrails as strong as in English?

The last item is the most dangerous in practice, and is probably what Meta means: guardrails are tested and tuned in English, and may be weaker in other languages. Anyone building a public Arabic product should add a protection layer of their own rather than relying on the model’s alone.


What “performance threshold” means in practice

Meta’s phrasing is precise and worth unpacking:

The part The practical meaning
“Performance thresholds” There is an internal number Arabic did not reach
“For safety” Protection is not guaranteed in your language
“And helpfulness” Quality is not guaranteed in your language
“May be able to output text” It will work, without commitment

The whole sentence is a carefully drafted disclaimer. And I respect its drafting because it is honest: the model will write Arabic, and the company does not warrant it.

The conclusion I want to leave: a language list is not a linguistic classification, it is the boundary of the warranty. Read it the way you read a device warranty: what is inside the list is covered, and what is outside is at your own risk.


The practical verdict

The eight-language list is a document, so use it. If your language is outside it, you are outside the scope of the commitment, however excellent the model.

And the licence permits you to fine-tune, so use that too. Meta wrote it explicitly. That is the best thing in this release for the region: a very strong foundation and explicit permission to build on it.

And do not wait for someone to measure your language. If the largest open model in the world has no Arabic figure two years on, waiting is not a plan.


Next in the series

On roughly the same day, Saudi Arabia’s Data and AI Authority published the paper for its ALLaM model. It contains a number worth reading twice: half its Arabic data is machine-translated. And in the model’s own card, a table placing a general Chinese model above the Saudi model on Arabic. The final article in this set.


Sources

  • Meta, Llama 3.1 model card: github.com/meta-llama/llama-models
  • Grattafiori et al., The Llama 3 Herd of Models, arXiv:2407.21783, July 2024
  • Global MMLU paper, arXiv:2412.03304, December 2024
  • AraLingBench, arXiv:2511.14295