Arabic AI Benchmarks Fractured—A New Leaderboard Tries to Unite 700+ Models
Curated by the Inblix editorial team
The rush to benchmark Arabic large language models has, ironically, created a mess. What started as a few narrow, author-specific demos has splintered into a landscape where at least four major leaderboards launched within seven months of 2024. The result is a community awash in scores but starved for clarity.
The original Open Arabic LLM Leaderboard (OALL), launched in May 2024 by 2A2I, TII, and HuggingFace, was meant to fix this. It offered 14 centralized benchmarks covering reading comprehension, sentiment analysis, and question answering. The community responded in force: over 46,000 visitors, 700 submitted models from 180 unique organizations, and 8 academic citations. It became, by traffic and engagement, the most prominent Arabic evaluation platform to date. But activity isn’t the same as utility.
The problem is that while OALL solved the resource barrier—users no longer needed to rent GPUs just to see how a model performed—it couldn’t solve the integrity problem. The first iteration still relied on a system where, as the organizers themselves note, there was “no robust mechanism to ensure those results were accurate.” When a leaderboard becomes popular, the incentive to game it rises proportionally. Meanwhile, competitors upped the ante: SDAIA’s Balsam Index flooded the field with 1,400 datasets and 50,000 questions, Scale’s SEAL lab locked its prompts behind private human-preference testing, and the AraGen Leaderboard shifted focus to generative tasks with a culturally-aware benchmark.
It’s a classic benchmark arms race. More tests, more private datasets, more evaluation dimensions. But for a developer trying to pick a model for a downstream application today, the fragmentation is paralyzing. Do you trust the public scores? Do you prioritize generative fluency or reading comprehension? The sheer volume of submitted models—over 70% of which are chat and fine-tuned variants, with half sitting under 7B parameters—suggests a community experimenting wildly but without a shared yardstick.
This is the exact tension a second version of the leaderboard must resolve. The original OALL proved the demand is there; Arabic, one of the world’s most spoken languages with relatively limited web presence, clearly has a vibrant open-source AI scene. What it lacked was a verification layer that could make those comparisons trustworthy. The next iteration isn’t just a nice-to-have. It’s the difference between informed model selection and performance theater.
💡 Key Takeaways
- Over 700 Arabic LLMs were submitted to the first Open Arabic LLM Leaderboard from 180+ organizations, but the platform had no mechanism to verify the accuracy of self-reported scores.
- The Arabic benchmarking space fragmented rapidly in 2024, with four major leaderboards launching in seven months—each using different evaluation methods from private human-preference tests to 1,400-dataset suites.
- Fine-tuned and chat models dominate submissions at over 70%, while models under 7B parameters constitute more than half the field, signaling a community focused on lightweight, application-specific deployments rather than massive foundational training.
- A new leaderboard version must solve the trust gap: centralized evaluation proved demand, but without verification, high traffic and citation counts don't translate into reliable model comparison.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.