Karpathy spent $10 and got Claude Opus 5 to build a Lord of the Rings 3D world
The Decoder · Aug 3, 2026 · 2 min read
Ten bucks and a single paragraph from Tolkien. That’s what Andrej Karpathy fed Claude Opus 5, and what he got back was...
9 articles
Explore our coverage of AI benchmarks — 9 curated articles, summaries, and related resources from the Inblix archive.
The Decoder · Aug 3, 2026 · 2 min read
Ten bucks and a single paragraph from Tolkien. That’s what Andrej Karpathy fed Claude Opus 5, and what he got back was...
The Verge AI · Aug 3, 2026 · 3 min read
Alibaba just dropped Qwen3.8-Max, a 2.4-trillion-parameter model that doesn't just narrow the gap with US frontier labs...
The Decoder · Jul 24, 2026 · 2 min read
Anthropic dropped a new flagship model that reshapes the price-performance conversation. Claude Opus 5 doesn't just inc...
The Decoder · Jul 21, 2026 · 2 min read
Google dropped three new Gemini models on Tuesday, but the launch felt less like a power move and more like a stall tac...
OpenAI Blog · Jul 15, 2026 · 2 min read
How good are AI agents at doing the actual job of a machine learning engineer? Not just writing snippets, but wrangling...
OpenAI Blog · Jul 11, 2026 · 1 min read
OpenAI dropped a new benchmark called FrontierScience to see if AI models can actually do expert-level scientific reaso...
OpenAI Blog · Jul 10, 2026 · 1 min read
SWE-bench Verified, once the gold standard for measuring AI models' autonomous software engineering skills, is now effe...
The Decoder · Jul 9, 2026 · 1 min read
OpenAI just dropped a bombshell on the AI coding world: the popular SWE-Bench Pro test, which measures how well AI mode...
The Decoder · Jul 8, 2026 · 1 min read
Anthropic's Claude Fable 5 is the new king of AI benchmarks, crushing eight industry-specific indices from finance to m...