AI Pulse by Inblix

Up to 40% of Arabic AI Benchmarks Are Broken, QIMMA Audit Finds

Hugging Face Blog · Apr 21, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Up to 40% of Arabic AI Benchmarks Are Broken, QIMMA Audit Finds

If you’ve been tracking Arabic LLM evaluation, you’ve probably noticed a growing tension: the number of benchmarks and leaderboards is expanding rapidly, but are we actually measuring what we think we’re measuring? A new open-source initiative suggests the answer is a firm no. QIMMA, which means ‘summit’ in Arabic, ran a rigorous quality audit on over 52,000 samples from 14 established Arabic benchmarks before evaluating any models. The results are sobering. As many as 40 percent of samples in some widely-used benchmarks were flagged or eliminated for containing systematic quality issues. We’re not talking about edge cases. The problems included factually wrong gold answers, corrupt text, encoding errors, and answers that didn’t comply with the evaluation protocol. One benchmark had gold indices that simply didn’t match any of the answer choices. Another contained stereotype-reinforcing content that flattened diverse Arab communities into monolithic generalizations. To catch these, QIMMA’s pipeline first used two strong LLMs—Qwen3-235B and DeepSeek-V3—to independently score each sample against a 10-point rubric. If either model scored a sample below 7 out of 10, it got a ticket to human review. Native Arabic speakers then made the final call, weighing dialectal nuance and cultural context that automated systems might miss. What’s left is a clean evaluation suite that is 99 percent native Arabic content, spans seven domains from healthcare to software development, and includes the first Arabic leaderboard with code evaluation. The team adapted HumanEval+ and MBPP+ with Arabic problem statements. The implication is uncomfortable but clear: a good chunk of the Arabic NLP leaderboard you’re looking at today is likely contaminated by bad data. QIMMA offers a cleaner starting point, but it also raises a sharp question about every other benchmark we take for granted.

💡 Key Takeaways

  1. A multi-stage validation pipeline found that established Arabic benchmarks contain systematic flaws, including factually incorrect answers and misaligned gold labels, leading to the elimination of up to 40% of samples in some cases.
  2. QIMMA is the first Arabic evaluation suite to combine native content, code evaluation via Arabic-adapted HumanEval+ and MBPP+, and public per-sample model outputs for full reproducibility.
  3. The audit revealed cultural sensitivity problems, with some benchmarks reinforcing stereotypes or treating diverse Arab cultures as monolithic, an issue automated checks frequently miss without native speaker review.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles