Product
OpenAI Finds SWE-bench Is Broken—So They're Fixing It
OpenAI Blog · Jul 16, 2026 · 2 min read
SWE-bench has been the gold standard for proving your AI can code like a real engineer. Top models barely scraped 20% o...
3 articles
Explore our coverage of evaluation — 3 curated articles, summaries, and related resources from the Inblix archive.
OpenAI Blog · Jul 16, 2026 · 2 min read
SWE-bench has been the gold standard for proving your AI can code like a real engineer. Top models barely scraped 20% o...
Hugging Face Blog · May 6, 2026 · 2 min read
The Open ASR Leaderboard just got its first taste of private test data — and it's a direct shot at the subtle art of 'b...
Hugging Face Blog · Jun 6, 2025 · 2 min read
Evaluating AI that can actually use a computer like a person — by looking at the screen — remains surprisingly hard to...