93 on Arabic, and One Question Exposes the Rest: Gemini 3 Pro and Cultural Reasoning

An overhead view of a laden table with hands offering and refusing at once

Ehab Saleh | techkahwa.net Model released: November 18, 2025 | Published: December 2, 2025 The numbers first Metric Result What it measures Source Global-MMLU-Lite Arabic 93, the highest on the index General Arabic knowledge and reasoning Artificial Analysis Arabic index Artificial Analysis Intelligence Index First, three points above GPT-5.1 Composite general intelligence Artificial Analysis, November … Read more

“Bukra ʿindi meeting”: Testing GPT-5.1 on the Arabic We Actually Write

Two rivers of light braiding together without their colours mixing

Ehab Saleh | techkahwa.net Model released: November 12, 2025 | Published: November 26, 2025 The numbers first Metric Result What it measures Source HELM Arabic Among the closed models evaluated on seven native Arabic benchmarks Independent Arabic performance Stanford CRFM and Arabic.AI, December 2025 HELM Arabic leader Arabic.AI LLM-X Arabic performance on seven native benchmarks … Read more

4.74 in MSA and 2.73 in Levantine: The Number That Sums Up the Arabic Evaluation Crisis

A polished monument beside the same monument drawn as a faint wireframe

Ehab Saleh | techkahwa.net Model released: August 2025 | Published: September 14, 2025 The numbers first An independent evaluation of ALLaM 34B through the HUMAIN Chat interface, judged by three separate frontier models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4), over 23 prompts repeated five times each for 115 total responses: Category Score out of … Read more

A Dialect Is Not an Identity: The Stereotype Trap on Grok 4

A crowd of identical silhouettes with one singled out by a harsh spotlight

Ehab Saleh | techkahwa.net Model released: July 10, 2025 | Published: July 24, 2025 The numbers first Metric Result What it measures Source Artificial Analysis Intelligence Index 22 on the index scale at time of measurement Composite general intelligence Artificial Analysis Context window 256K tokens Text processed in a single pass Artificial Analysis Price $3 … Read more

Fluency Is Not Understanding: Qwen3 and the Best Arabic Score Among Open Models

An exquisitely cut glass vessel that is completely hollow inside

Ehab Saleh | techkahwa.net Model released: April 28, 2025 | Published: May 12, 2025 The numbers first Metric Result What it measures Source HELM Arabic, best open-weights model Qwen3 235B, mean 0.786 Seven native Arabic benchmarks Stanford CRFM and Arabic.AI, December 2025 HELM Arabic, open models in the top ten Four, including Qwen3 235B and … Read more

MSA Leakage: How Llama 4 Obeys Your Dialect Request and Violates It at the Same Time

A colourful mosaic wall being swallowed from below by grey concrete

Ehab Saleh | techkahwa.net Model released: April 5, 2025 | Published: April 19, 2025 The numbers first Metric Result What it measures Source HELM Arabic Llama 4 Maverick among the top ten open-weights models Seven native Arabic benchmarks Stanford CRFM and Arabic.AI, December 2025 OALL leaderboard, chat models Llama3.3-70B-Instruct in first place Native Arabic benchmarks … Read more

Two Years of Arabic: What Actually Changed Since Claude 3.5 Sonnet and Gemini 2.5 Pro

Layered geological strata with one core sample drilled through all of them

Baseline releases: October 2024 and March 25, 2025 | Published: April 8, 2025 Ehab Saleh | techkahwa.net The numbers first What changed Before After Source OALL leaderboard Version one, partly built on machine-translated benchmarks Version two dropped the translations for native benchmarks: ArabicMMLU, MadinahQA, AraTrust, ALRAGE OALL v2 Arabic speech measurement No multidialectal shared task … Read more

Which Language Does the Model Think In Before Answering You in Arabic? The DeepSeek R1 Case

A transparent head full of gears crossed by a single unchanged thread of light

Ehab Saleh | techkahwa.net Model released: January 20, 2025 | Published: February 3, 2025 The numbers first Metric Result What it measures Source XReasoning, thinking-language match rate 46.3% on AIME and 42.3% on GPQA for the largest R1-distilled model Does the model think in the language it was asked to? arXiv 2505.22888 XReasoning, after forcing … Read more

Half of Its Arabic Is Machine-Translated: Reading ALLaM and What “National Model” Means

ARAB-LENS article illustration 38

Ehab Saleh | techkahwa.net Paper released: July 22, 2024 | Published: August 12, 2024 The numbers first Metric Result What it measures Source Total Arabic curated 540 billion tokens The size of the Arabic corpus ALLaM paper, arXiv:2407.15390 Of which natural Arabic 270 billion Half the total Same paper, quoted verbatim Of which machine-translated Arabic … Read more

Eight Supported Languages, and Arabic Is Not One: Llama 3.1 at 405 Billion

ARAB-LENS article illustration 37

Ehab Saleh | techkahwa.net Model released: July 23, 2024 | Published: August 6, 2024 The numbers first Metric Result What it measures Source Officially supported languages Eight: English, German, French, Italian, Portuguese, Hindi, Spanish, Thai Meta’s declared commitment Llama 3.1 model card Arabic among them No Whether your language is on the list Same card … Read more