The Vowel That Flips the Meaning: DeepSeek V4 Pro on Arabic Morphology and Syntax

Reading Time: 13 min
18
Ehab Saleh | techkahwa.net Model released: April 24, 2026 | Published: May 8, 2026

The numbers first

MetricResultWhat it measuresSource
Artificial Analysis Intelligence Index52, second among open weights reasoning modelsComposite general intelligenceArtificial Analysis, 2026
GDPval-AA1554, first among open weights on agentic workPerformance on real work tasksArtificial Analysis
Architecture1.6T total parameters, 49B activeSize and costArtificial Analysis
Context window1M tokens, an 8x expansion over V3.2Processing capacityArtificial Analysis
Price$1.74 / $3.48 per million tokensCostArtificial Analysis
LicenseMIT, open weightsCan be self-hostedArtificial Analysis
HELM Arabic, predecessor V3.1Among the top ten open-weights modelsArabic performance on seven native benchmarksStanford CRFM and Arabic.AI, December 2025
Any published diacritics or morphology metricNone for this releaseAccuracy of morphology and vowellingSurvey of public sources

The model ships with open weights under an MIT license, and that is not a legal footnote, it is why I am writing about it. Any Arab university or company can run it locally and measure it themselves, without anyone’s permission and without sending data abroad. Which gives the question “how is its Arabic” a different practical weight here.


The forgotten axis in Arabic evaluation

When people discuss Arabic capability, the conversation goes to knowledge, culture, and dialect. The structure itself, morphology, syntax, orthography, diacritics, gets treated as solved. It is not.

Arabic is written without short vowels in everyday use. Which means the reader, human or model, infers the vowels from context:

د ر س
→
darasa
(he studied)
→
darrasa
(he taught)
→
dars
(a lesson, noun)
→
durisa
(it was studied, passive)

Four meanings in three letters. A model that gets this wrong is not making a diacritics error, it is swapping the subject and object of the sentence. And this is not theoretical: the NADI 2025 diacritic restoration task for spoken dialects recorded 55 WER and 13 CER in the best participating systems, meaning more than half the words came out incorrectly vowelled among the strongest competitors.


Model card

FieldDetail
NameDeepSeek V4 Pro
LabDeepSeek
ReleasedApril 24, 2026, alongside V4 Flash
TypeHybrid thinking and non-thinking model
InputsText only
LicenseMIT
Axis under testARAB-LING: morphology, syntax, orthography, diacritics

The test: ten items that expose structure

These items are built from the errors I actually see in Arabic model output, not from a grammar textbook. Each one exposes a different layer.

1. Non-human plural agreement. Ask it to describe books, cities, and cars. Correct Arabic takes feminine singular agreement for non-human plurals: al-kutub jadida, not al-kutub judud. This is the single most common grammatical error I see in machine-written Arabic, and the easiest test for a model that is mapping English structure onto Arabic words.

2. Agreement by word order. Ask for the same meaning twice, once verb-first and once subject-first. Correct: dhahaba al-tullab when the verb leads (singular verb), al-tullabu dhahabu when the subject leads (full agreement). The error: dhahabu al-tullab.

3. The construct state (idafa). Ask for “the student’s new book” with “new” describing the book. The model must know the first noun takes no definite article and no tanwin, and that no adjective may sit between the two parts. Plenty of strong models fail this one.

4. Number agreement. Ask for three, seven, eleven and twenty with both masculine and feminine counted nouns. Correct: thalathatu rijal and thalathu fatayat with reversed gender polarity, and ʿishruna kitaban with a singular accusative.

5. Meaning-distinguishing diacritics. Give it a sentence where context forces the doubled form, and ask for full vowelling with reasoning.

6. Hamza placement. Ask it to write ra’is, mas’ul, sa’ala, shay’, mubtadi’un. Errors here reveal a model trained on poorly spelled Arabic text, which is an indirect signal about the quality of its Arabic training data.

7. Ta marbuta versus ha. Ask it to distinguish madrasatuhu from madrasatun in one context.

8. Alif maqsura versus ya. ʿala the preposition against ʿalayya, ila against ilayya.

9. Arabic punctuation. Ask for a paragraph using Arabic punctuation marks: the Arabic comma, semicolon and question mark. A model that emits Latin commas inside Arabic text is telling you it treats Arabic as an encoding layer over English.

10. Reverse proofreading. Give it a paragraph with five deliberate errors and ask it to find, fix and justify each. This is the strongest item on the list, because correct generation can be memorization, while catching an error requires knowing the rule.

Difficulty ordering as I observe it in practice
Hamza and spelling
most models are fine
Non-human agreement
inconsistent
Number agreement
weak
Construct state
weak
Distinguishing vowels
weakest
Reverse proofreading
weakest

Why the published metric misleads you

OALL version two includes a benchmark called MadinahQA, dedicated to Arabic language and grammar knowledge, and it also appears among the seven benchmarks in HELM Arabic. The good news is that it exists at all. The less good news is that, like most of its neighbors, it is multiple choice.

The distance between “pick the correct answer about a grammatical rule” and “write me a grammatically sound paragraph” is the distance between a student who memorized the rule and a writer who commands it. A model can answer a question about non-human plural agreement correctly, then write al-kutub judud in the next paragraph. I have watched it happen more than once.

So ARAB-LENS carries a rule here: never accept a grammar score from multiple choice alone. The real measurement is an error rate over five hundred words of generated text, counted by hand across six categories: agreement, construct state, numbers, hamza, ta marbuta, and punctuation.


The practical verdict

An open model under an MIT license offers something the closed ones cannot: you can fine-tune it on clean Arabic yourself. That is the most important recommendation here.

  1. Do not rely on a general model for Arabic proofreading. Make proofreading a separate layer with explicit rules, then let the model propose and the rule decide.
  2. Measure errors per thousand words, not by impression. Six categories, a hand count over five samples, and you will know your model better than any leaderboard does.
  3. Invest in correctly spelled Arabic data if you intend to fine-tune. Hamza quality in the output follows hamza quality in the data, almost without exception.
  4. If your product is educational, diacritics are not a luxury. And with NADI’s current numbers, do not hand automatic vowelling to an end user without review.

Next in the series

To the other end of the spectrum: a model built for this language specifically. Falcon-H1-Arabic from the Technology Innovation Institute in Abu Dhabi, with the full OALL numbers and one question: does specializing in Arabic beat raw scale?


The benchmarks themselves are not clean

A finding published in April 2026 deserves to be known by anyone citing Arabic leaderboards. The QIMMA leaderboard from the Technology Innovation Institute built its methodology on a different idea: instead of aggregating existing benchmarks as they are, it first ran them through a quality validation pipeline, with multi-model automated review followed by human review, discarding broken samples before evaluation.

The result is revealing: ArabicMMLU, among the most cited benchmarks in Arabic evaluation, recorded a discard rate of 3.1% due to answer quality and text formatting problems.

Three percent looks small and is large in evaluation terms: the gaps between closely ranked models are often under three points. Which means the noise inside the benchmark can equal or exceed the difference whole rankings are built on.

QIMMA’s other numbers are worth citing: more than 52,000 samples across fourteen benchmarks and seven domains, 99% natively Arabic rather than translated, and it is the first Arabic leaderboard to add code generation. In its first results Qwen3.5-397B-A17B led with a mean of 68.06, followed by Karnak at 66.20 and Jais-2-70B-Chat at 65.81. And the leaderboard’s own observation matters: Arabic-specialized models lead on cultural and linguistic tasks, while multilingual models lead on code.

What Arabic localizers actually complain about

In Arabic localization and translation circles, the list of complaints about model output has been stable for years, in order of frequency:

  1. Ambiguity in undiacritized text, specifically confusing base and derived verb forms.
  2. Plural agreement errors, most famously treating non-human plurals as human.
  3. Literal translation of idioms.
  4. Weak code-switching, translating technical terms that are normally left in English.
  5. Weak specialist terminology in medicine, law and engineering.

That list maps almost item for item onto the test in this article, and the overlap is itself the argument: when what a localizer notices in daily work matches what a designed test exposes, you are looking at a real failure pattern rather than a coincidence.

The audit book: six categories with real error examples

These six categories are what I count in every generated Arabic text. The table gives the form each error actually takes, and its correction:

CategoryHow the error appears in outputCorrect formWhy it happens
Non-human plural agreemental-kutub judud, al-mudun kibaral-kutub jadida, al-mudun kabiraTransferring English plural structure into Arabic
Verb and subject orderdhahabu al-tullab ila al-saffdhahaba al-tullab, or al-tullabu dhahabuIgnoring how fronting changes agreement
Construct stateal-kitabu al-talibi al-jadidkitabu al-talibi al-jadidPutting the article on the first noun
Number agreementthalath rijal and thalathat fatayatthalathatu rijal and thalathu fatayatNot knowing the gender polarity rule
Hamzamas’ul written mas’ul with the wrong seat, shay’, mubtadi’unCorrect seats per the strongest-vowel ruleTraining on poorly spelled Arabic text
PunctuationAn Arabic sentence ending in ? or ,Ending in the Arabic ؟ or ،Treating Arabic as an encoding over English

The first row alone tells me more than half of what I need to know about a model. Anyone writing al-kutub judud is translating a structure. Anyone writing al-kutub jadida is writing Arabic.

Turning the count into a number

The method I use: five texts of five hundred words in one domain, so 2,500 words total, a manual count across the six categories, then multiply by four for a per-thousand-word rate.

A worked arithmetic illustration: if across those 2,500 words you counted seven agreement errors, three construct-state errors and five hamza errors, your rates are 2.8, 1.2 and 2.0 per thousand words, and your combined linguistic error rate is six per thousand.

The number alone means nothing until you repeat it. Its value appears when you measure the same model six months later, or compare two models on identical texts. Then you hold what no public leaderboard gives you: a curve specific to the kind of writing your product actually needs.

And a practical observation from experience: errors do not distribute evenly. Hamza and punctuation improve quickly with each generation, while agreement and construct state stay stubborn, because they are structural rather than orthographic. So if you find a model with few spelling errors and many structural ones, you are looking at a model that read a lot of Arabic without learning its rules, which is worse than the reverse.


From the proofreading notebook: five sentences I use every time

These five are short and deliberately dull, and they require no knowledge of the model and no tooling. I send them as they are and read the reply with a proofreader’s eye:

Sentence one: Write a paragraph about three old Syrian cities. What I watch: al-mudun al-qadima (feminine singular agreement) rather than al-mudun al-qudama, and thalath mudun with reversed gender rather than thalatha mudun.

Sentence two: Describe five books you bought. What I watch: khamsat kutub with the masculine numeral because kitab is masculine, and al-kutub mufida rather than al-kutub mufidun.

Sentence three: Write a sentence containing the student’s name and an adjective for the book. What I watch: kitabu al-talibi al-jadid, not al-kitabu al-talibi al-jadid, and no adjective inserted between the two parts of the construct.

Sentence four: Write a question and then an answer using Arabic punctuation. What I watch: the Arabic ؟ and ، rather than the Latin ? and ,. This item quickly reveals whether the model treats Arabic as a language or as an encoding.

Sentence five: Correct this paragraph and explain each correction: “al-tullab alladhina najahu fi al-imtihan kanu khamsat talibat, wa qad qara’a al-mudarrisa asma’ihim ʿala al-mala’.” There are five deliberate errors here: a wrong hamza on al-imtihan, broken gender polarity in khamsat talibat, qara’a where qara’at is required, asma’ihim where the accusative asma’ahum is required, and al-mala’ misspelled. A model that finds three and explains them is good, one that finds five and explains them is excellent, and one that silently rewrites the paragraph without explanation has not been tested at all.

The effect of missing diacritics, in one example

This sentence, unvowelled, carries two contradictory readings:

ʿallama al-mudir al-muwazzaf al-jadid.

Did the manager teach the new employee, or was the manager the one taught? And does ʿallama mean taught or marked? An Arabic speaker resolves it instantly from context, while a model sometimes resolves it by the statistically likelier reading rather than by the context.

So every evaluation I run carries one item of this kind: an ambiguous sentence with context that settles it, and a question about the correct reading and why. Failure here is not a diacritics failure, it is a comprehension failure, and that is what makes it matter.

Sources

  • Artificial Analysis, DeepSeek V4 Pro and V4 Flash: artificialanalysis.ai/articles/deepseek-is-back-among-the-leading-open-weights-models-with-v4-pro-and-v4-flash
  • HELM Arabic, Stanford CRFM with Arabic.AI: crfm.stanford.edu/2025/12/18/helm-arabic.html
  • Open Arabic LLM Leaderboard v2: huggingface.co/blog/leaderboard-arabic-v2
  • NADI 2025, diacritic restoration for spoken dialects: arxiv.org/abs/2509.02038
  • QIMMA Arabic leaderboard and its sample validation pipeline: huggingface.co/blog/tiiuae/qimma-arabic-leaderboard
  • Practitioner test of eight models on Arabic translation: localazy.com/blog/ai-8-llm-arabic-models-tested-to-translate