AI Pulse by Inblix

Topic: evaluation

3 articles

Explore our coverage of evaluation — 3 curated articles, summaries, and related resources from the Inblix archive.

OpenAI Finds SWE-bench Is Broken—So They're Fixing It — Inblix summary
Product

OpenAI Finds SWE-bench Is Broken—So They're Fixing It

OpenAI Blog · Jul 16, 2026 · 2 min read

SWE-bench has been the gold standard for proving your AI can code like a real engineer. Top models barely scraped 20% o...

Hugging Face locks down new speech datasets to stop AI benchmark gaming — Inblix summary
Research

Hugging Face locks down new speech datasets to stop AI benchmark gaming

Hugging Face Blog · May 6, 2026 · 2 min read

The Open ASR Leaderboard just got its first taste of private test data — and it's a direct shot at the subtle art of 'b...

Hugging Face's ScreenSuite stress-tests 5 AI models on pure vision — no DOM cheating — Inblix summary
Research

Hugging Face's ScreenSuite stress-tests 5 AI models on pure vision — no DOM cheating

Hugging Face Blog · Jun 6, 2025 · 2 min read

Evaluating AI that can actually use a computer like a person — by looking at the screen — remains surprisingly hard to...