Arabic AI models flunk harder benchmarks: Top score plunges from 83% to 70%
Curated by the Inblix editorial team
The rug got pulled out from under Arabic large language models this week. MBZUAI and the Arabic-Leaderboards team tightened their evaluation gauntlet with the AraGen-03-25 release, and even the best model took a beating. The previous champion, OpenAI’s o1-2024-12-17, saw its score crater from a comfortable 82.67% down to a humbling 70.25%. That’s a 12-point drop just by swapping out the private test set and clarifying the judge’s instructions.
It’s a brutal recalibration that exposes just how much benchmark gaming was going on. The new dataset expands from 279 to 340 question-answer pairs, still heavily skewed toward question answering with around 200 pairs. But it’s not just about volume. The team, which is also releasing the old AraGen-12-24 test set and all model responses for public scrutiny, cranked up the difficulty on the blind test. The result? The clump of models that were previously jostling for second place in the 70-78% range got smooshed down into a 51-57% band. That’s not a minor correction; it’s a signal that those models were coasting on memorization rather than robust reasoning, particularly in Arabic.
“This highlights that the updated AraGen benchmark is more challenging and better highlights true performance gaps,” the team noted, with classic academic understatement. There’s a mystery in the data too: GPT-4o (version 2024-08-06) suddenly jumped up in the rankings. The researchers are scratching their heads over that one, calling it “under investigation.” A weird prompt sensitivity? A fluke? It’s a reminder that even evaluation science is messy. The overall ranking structure held steady though, with Claude-3.5-Sonnet as the judge and a refined system prompt designed to guide weaker evaluator models. The fact that prompt changes alone didn’t cause chaos is good news for reproducibility.
What’s genuinely new here is the arrival of the first public benchmark for Arabic instruction-following, Arabic IFEval. This isn’t just about knowing facts; it tests if a model can actually follow constraints like “write in a specific format” or “use a certain word.” It’s a harder, more useful test of whether an AI assistant is actually helpful. For anyone building on Arabic AI, the message from this unified leaderboard space is clear: the vibes-based era of evaluation is over. If your model’s scores didn’t just get halved, you might actually have something real.
💡 Key Takeaways
- OpenAI's o1 model remains the most reliable for Arabic, but its score plummeted from 82.67% to 70.25% on the harder AraGen-03-25 benchmark, revealing previous tests were too easy.
- Scores for second-tier models collapsed from the 70-78% range into the 51-57% band, suggesting their prior high performance relied on dataset-specific patterns rather than robust language understanding.
- The launch of the Arabic IFEval benchmark provides the first standardized tool to measure if models can follow granular instructions in Arabic, moving beyond simple accuracy tests.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.