Mira Murati's Inkling Small matches bigger sibling on 12B parameters, not 276B
Curated by the Inblix editorial team
Mira Murati’s Thinking Machines just dropped Inkling Small, and the numbers tell a story that flips the dominant AI narrative on its head. The open-weights reasoning model scores 40 on Artificial Analysis’s Intelligence Index — just one point shy of its bigger sibling, Inkling, which scored 41. But here’s the kicker: it does this with less than a third of the parameters. We’re talking 276 billion total parameters but only 12 billion active at any given time. According to AA, no open model of equal or smaller size posts a higher score.
The model actually outperforms its larger counterpart on several benchmarks that developers care about. Think Humanity’s Last Exam, where it hit 32% compared to Inkling’s 30%, and GPQA Diamond, where it scored 89% versus 87%. It stumbles a bit on agent-based tasks and raw factual knowledge, but the trade-off in token efficiency is substantial. Inkling Small averages just 24,000 output tokens per task. For context, that’s nearly half of what Deepseek V4 Flash consumes at 45,000 tokens and less than a third of GPT-5.4 mini’s 78,000-token appetite.
Thinking Machines is shipping this thing under Apache 2.0, with weights already live on Hugging Face. It handles text, image, and speech inputs, sports a 256,000-token context window, and users can fine-tune it right in the browser through something called Tinker Playground. The positioning here is deliberate: Murati’s team isn’t chasing the biggest model crown. They’re selling a foundation that others can build on with their own data.
Some observers are calling this the next frontier in AI, and for once that might not be hyperbole. A model that delivers near-identical reasoning chops at a fraction of the computational cost changes the calculus for anyone deploying these systems at scale. The question isn’t whether smaller, more efficient models will catch on — it’s whether labs still burning cash on massive monolithic models will take the hint. Murati clearly already has.
💡 Key Takeaways
- Inkling Small achieves a 40 on the Intelligence Index with only 12B active parameters, matching its larger sibling's 41 while using far less compute.
- The model averages 24K output tokens per task, making it roughly 3x more efficient than GPT-5.4 mini for equivalent reasoning benchmarks.
- Released under Apache 2.0 with browser-based fine-tuning via Tinker Playground, the model is explicitly built as a customizable foundation rather than a monolithic product.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.