AI Pulse by Inblix

Topic: AI benchmarks

9 articles

Explore our coverage of AI benchmarks — 9 curated articles, summaries, and related resources from the Inblix archive.

Karpathy spent $10 and got Claude Opus 5 to build a Lord of the Rings 3D world — Inblix summary
AI News

Karpathy spent $10 and got Claude Opus 5 to build a Lord of the Rings 3D world

The Decoder · Aug 3, 2026 · 2 min read

Ten bucks and a single paragraph from Tolkien. That’s what Andrej Karpathy fed Claude Opus 5, and what he got back was...

Alibaba's 2.4-trillion-parameter Qwen3.8-Max dethrones Claude Fable 5 on key benchmarks — Inblix summary
AI News

Alibaba's 2.4-trillion-parameter Qwen3.8-Max dethrones Claude Fable 5 on key benchmarks

The Verge AI · Aug 3, 2026 · 3 min read

Alibaba just dropped Qwen3.8-Max, a 2.4-trillion-parameter model that doesn't just narrow the gap with US frontier labs...

Claude Opus 5 just scored 4x higher than GPT-5.6 Sol on novel problem-solving — Inblix summary
AI News

Claude Opus 5 just scored 4x higher than GPT-5.6 Sol on novel problem-solving

The Decoder · Jul 24, 2026 · 2 min read

Anthropic dropped a new flagship model that reshapes the price-performance conversation. Claude Opus 5 doesn't just inc...

Google ships 3 budget Gemini Flash models while its missing Pro cedes ground to rivals — Inblix summary
AI News

Google ships 3 budget Gemini Flash models while its missing Pro cedes ground to rivals

The Decoder · Jul 21, 2026 · 2 min read

Google dropped three new Gemini models on Tuesday, but the launch felt less like a power move and more like a stall tac...

OpenAI's o1-preview can snag Kaggle bronze, new MLE-bench reveals — Inblix summary
Product

OpenAI's o1-preview can snag Kaggle bronze, new MLE-bench reveals

OpenAI Blog · Jul 15, 2026 · 2 min read

How good are AI agents at doing the actual job of a machine learning engineer? Not just writing snippets, but wrangling...

Product

FrontierScience tests how well AI can do real science

OpenAI Blog · Jul 11, 2026 · 1 min read

OpenAI dropped a new benchmark called FrontierScience to see if AI models can actually do expert-level scientific reaso...

Product

SWE-bench Verified no longer measures real coding ability

OpenAI Blog · Jul 10, 2026 · 1 min read

SWE-bench Verified, once the gold standard for measuring AI models' autonomous software engineering skills, is now effe...

AI Coding Test SWE-Bench Pro Has 30% Flawed Tasks — Inblix summary
AI News

AI Coding Test SWE-Bench Pro Has 30% Flawed Tasks

The Decoder · Jul 9, 2026 · 1 min read

OpenAI just dropped a bombshell on the AI coding world: the popular SWE-Bench Pro test, which measures how well AI mode...

Claude Fable 5 tops industry benchmarks with huge price gap — Inblix summary
AI News

Claude Fable 5 tops industry benchmarks with huge price gap

The Decoder · Jul 8, 2026 · 1 min read

Anthropic's Claude Fable 5 is the new king of AI benchmarks, crushing eight industry-specific indices from finance to m...