AI Pulse by Inblix

OpenAI admits its own benchmarks are useless for 80% of the world

OpenAI Blog · Jul 12, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI admits its own benchmarks are useless for 80% of the world

OpenAI just admitted something that anyone working in multilingual AI has known for years: the benchmarks we use to measure language understanding are broken. The company’s new IndQA benchmark is a direct response to the fact that existing tests like MMMLU are effectively saturated—top models cluster near perfect scores, making them useless for tracking actual progress. That’s a problem when roughly 80 percent of the world doesn’t speak English as their primary language.

So why India? The math is straightforward. About a billion people there use something other than English day to day, the country has 22 official languages, and it happens to be ChatGPT’s second largest market. The team partnered with 261 domain experts across India to build a set of 2,278 questions spanning 12 languages and 10 cultural domains—everything from Architecture to Religion to Sports. What makes IndQA different is that it doesn’t care about translation accuracy or multiple-choice guessing. It tests whether a model can actually reason about culturally specific topics, like a Bengali literature question or a Hinglish food query, where context is everything.

Each question comes with a rubric written by those same domain experts, spelling out exactly what an ideal answer should include. A model-based grader then checks responses against that rubric, assigning weighted points per criterion. The questions themselves were adversarially filtered: they only kept the ones where a majority of OpenAI’s own strongest models—GPT‑4o, o3, GPT‑4.5, and even GPT‑5 post-launch—failed to produce acceptable answers. That’s a brutal filter, and it means IndQA has plenty of headroom.

The results are, predictably, humbling. OpenAI’s models have improved on Indian languages over time, but the company is upfront that “substantial room for improvement” remains. They also caution against using IndQA as a cross-language leaderboard, since the questions aren’t identical across languages. Instead, it’s a yardstick for measuring a single model family’s improvement over time. The broader signal here is clear: if AI is going to work for everyone, we need evaluations that capture the messy, culturally embedded ways people actually use language—not just how well a model parses a translated textbook paragraph.

💡 Key Takeaways

  1. Existing multilingual benchmarks like MMMLU are saturated, with top models scoring so high that the tests no longer reveal meaningful differences in capability.
  2. IndQA was built adversarially against OpenAI’s strongest models, meaning every question in the set is one that GPT‑4o, o3, GPT‑4.5, and GPT‑5 could not answer acceptably.
  3. OpenAI explicitly warns that IndQA is not a language leaderboard because questions aren’t identical across languages, limiting cross-language comparisons.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles