Five Languages, and Arabic Is Not One of Them: Reading Mixtral 8x7B

Reading Time: 9 min
18
Ehab Saleh | techkahwa.net Model released: December 11, 2023 | Published: December 25, 2023

The numbers first

Metric Result What it measures Source
Declared languages Five: English, French, Italian, German, Spanish What the lab says about its own model Mistral announcement, December 11, 2023
Arabic mentioned in the announcement Not present Whether it was mentioned at all Same announcement
Arabic mentioned in the paper abstract Not present The technical document arXiv:2401.04088
Architecture Mixture of experts, 8 experts with 2 routed per token How compute is allocated Announcement and paper
Total parameters About 47 billion Nominal size Paper
Active parameters per token About 13 billion Actual cost Paper
Published Arabic result for the official model None exists I searched the OALL leaderboard and HELM Arabic OALL results dataset, HELM Arabic
What is on the OALL leaderboard instead Community fine-tunes only They do not represent the original model OALL results dataset

The precise formulation I hold to: Mistral’s announcement lists five languages, and Arabic is not among them. The company did not say the model does not support Arabic, and I will not say what they did not.


The language list, as published

The languages Mistral declared Mixtral masters
English
declared
French
declared
German
declared
Spanish
declared
Italian
declared
Arabic
not mentioned

Model card

Item Detail
Model Mixtral 8x7B
Lab Mistral AI, France
Release date December 11, 2023
Technical paper arXiv:2401.04088, January 8, 2024
Architecture Sparse mixture of experts
Licence Apache 2.0, fully open
Standing at the time The best open model on most English leaderboards
Arabic Not mentioned in any official document

The ARAB-LENS reading: the difference between being forgotten and being excluded

This article is about a fine distinction that matters.

With Llama 2 we saw an omission: a data table Arabic never reached the threshold of, and a general warning that the model may not suit other languages. No list, no declared decision.

With Mixtral we see something else: an explicit list. Five languages written by name. Which means somebody sat down and decided: these are the languages we commit to, and these are the ones we do not. Arabic is outside the list by decision, not by accident.

And I will say something that may surprise you: I prefer this to the omission.

Why? Because an explicit list is a document you can lean on. When you go to your manager and say “we cannot use this model for our Arabic product”, you have a written line from the company itself. An omission, by contrast, leaves a door open for anyone who wants to say “but it writes Arabic, I tried it”. And that door is the one through which dozens of failed projects I have watched walked in.

The second point, technical and important: Mixtral is a mixture-of-experts model. Eight experts, two routed per token. In practice that means the model has large capacity but spends it selectively.

Why does that matter for Arabic? Because a mixture-of-experts architecture specialises. The experts distribute themselves across patterns in the data during training. If there is not enough Arabic in the data, no expert will specialise in it, and Arabic will be handled by general experts designed for something else. And that is worse than a small dense model, because a dense model at least passes everything through all of its parameters.

Put more simply: in a mixture-of-experts model, a language you did not train on is not merely underserved, it is routed to the wrong place.

The third point, a direct methodological warning: when I went looking for an Arabic result for Mixtral 8x7B, what I found on the Open Arabic LLM Leaderboard were entries named dolphin-mixtral, Nous-Hermes-2-Mixtral and Smaug-Mixtral. These are not Mixtral. They are versions modified by individuals and communities. And the score of a modified version is not attributable to the original model, exactly as the performance of a tuned car is not attributable to the factory.

I see this error repeated often in Arabic coverage: a number is taken from a leaderboard, attributed to the company, and a sentence is built on it. Always check the model owner’s name on the leaderboard before you quote its number.


From the test notebook: how to test a model that promises you no Arabic

When a model sits outside your language’s list, what you need is not a score but boundaries. This is how I draw them.

Item one: where does forced code switching begin?

Ask for Arabic at five levels of difficulty and record when it escapes into English:

Level The request Expected behaviour
1 “Write a greeting sentence in Arabic” Succeeds
2 “Summarise this Arabic passage for me” Usually succeeds
3 “Explain how encryption works, in Arabic” Starts mixing English in
4 “Write a 500 word article in Arabic” Degrades after the second paragraph
5 “Convert this passage into Levantine” Fails or declines

Record the level. That number, not any benchmark, is what tells you where you can use it.

Item two: the four word test.

Four very simple Arabic words, each a different trap:

The word The trap The expected error
mish Egyptian and Levantine negation Read as a noun or ignored
leish Dialectal interrogative Confused with the Standard laysa
hallaʾ Levantine time adverb Mistranslated or dropped
bas Means both “only” and “but” Picks the wrong sense

Write: “mish ʿaaref leish hallaʾ, bas khalleena nʾajjelha.” and ask for an English translation. The correct one: “I do not know why now, but let us postpone it.” A model that renders bas as “only” has inverted the sentence.

Item three: the stability test.

Ask the same Arabic question five times in five separate conversations. A model weak in Arabic gives you five answers of different quality: one excellent, two mediocre, one in English, one mixed. Variance is the signal that Arabic is not a stable capability but a statistical accident.

In my view this item matters more than any score, because a product cannot tolerate variance. A model that gives you 70% every time is better for your product than one that gives 90% once and 40% the next.

Item four: who carries the liability?

A non-technical question that is the most important one. If the model errs in Arabic inside your product and a customer complains, what will you say? When the language is outside the declared list, the answer is that you alone are responsible. There is no vendor commitment to lean on. Make that part of your technical decision rather than part of your later surprises.

Item five: test output quality separately from input comprehension.

An important distinction that gets skipped: many models understand Arabic better than they write it. Test the two directions separately:

Direction The task What it measures
Comprehension “Summarise this Arabic passage in English” Understanding alone
Production “Express this English meaning in Arabic” Production alone
Both “Summarise this Arabic passage in Arabic” Both together

A model outside your language’s list usually passes the first, stumbles on the second and fails the third. And knowing which direction works opens uses you did not think were available: search, classification and extraction in Arabic are possible even with a model that writes poor Arabic.

Item six: the real cost test.

A model outside the list usually needs longer prompting and more attempts. Compute:

Tokens in the prompt + number of attempts until an acceptable answer

In my experience the gap between a supported and an unsupported model is not only quality but attempt count. And a model that is cheaper per token becomes more expensive per result if it needs three attempts instead of one.


A note on mixture-of-experts architecture and languages

A short technical section useful to anyone choosing open models.

Architecture How it handles a language that is scarce in its data
Dense Every parameter processes every token, so performance degrades smoothly
Mixture of experts Routing sends the token to experts that never specialised in it, so degradation is uneven

The practical consequence: mixture-of-experts models fluctuate more in unsupported languages. They may give you an excellent answer and then a poor one to nearly the same question.

That is a structural observation, not a fault in any particular model. But it means the stability test I mentioned above matters more with this architecture than with others.


The practical verdict

Read the language list before you read the leaderboard. That is the shortest piece of advice in this series and the biggest time saver in it.

If you must use a model outside the list, restrict it to tasks that do not require producing Arabic: classification, extraction, semantic search. Models do these in Arabic far better than they write it.

And never quote a leaderboard number without checking the model’s owner. Community fine-tunes are not the official model, and the gap between them can be twenty points in either direction.


Next in the series

In February 2024 an academic team published the first large Arabic exam built from real school questions, drawn from eight Arab countries, with not one machine-translated item. Fourteen thousand five hundred and seventy five questions. The next article is about ArabicMMLU, and about the numbers it exposed once translation was replaced by the original.


Sources

  • Mistral AI, Mixtral of experts, December 11, 2023: mistral.ai/news/mixtral-of-experts
  • Jiang et al., Mixtral of Experts, arXiv:2401.04088, January 2024
  • Open Arabic LLM Leaderboard results dataset: huggingface.co/datasets/OALL/results
  • HELM Arabic, Stanford Center for Research on Foundation Models: crfm.stanford.edu