AI Pulse by Inblix

Murati's Thinking Machines drops Inkling: US best, but trails China

The Decoder · Jul 16, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Murati's Thinking Machines drops Inkling: US best, but trails China

Thinking Machines Lab, the startup Mira Murati built after leaving OpenAI, just shipped its first model — and it’s a statement piece. Inkling is a 975-billion-parameter Mixture-of-Experts transformer that natively handles text, images, and audio. The weights are up on Hugging Face, no strings attached. But the real story here isn’t just the model’s size. It’s the positioning. The company, in its own announcement, is refreshingly blunt: “Inkling is not the strongest overall model available today.” They’re not selling a world-beater. They’re selling a platform for customization, with fine-tuning as the actual product.

On benchmarks, Inkling is a mixed bag that still manages to lead the US pack. Artificial Analysis rates it at 41 on their Intelligence Index, which puts it three points above the previous US open-weights leader, Nemotron 3 Ultra. It’s a genuine monster on agentic tasks, hitting an Elo of 1,238 on the GDPval-AA v2 benchmark, handily beating both Kimi K2.6 and DeepSeek v4 Flash max. It also shows off impressive token efficiency, using far fewer output tokens for the same tasks compared to those competitors. But then you look at the factual accuracy scores, and the picture gets complicated fast.

Here’s the brutal number: a 63 percent hallucination rate. Its accuracy sits at just 40 percent, earning it a score of +2 on the AA Omniscience benchmark. That’s better than some US models, but it’s a liability that will spook anyone considering this for work that requires ground-truth reliability. The cost doesn’t help either — $1.87 per million input tokens makes it pricier than capable Chinese open models like GLM-5.2 and DeepSeek v4. The model was trained on 45 trillion tokens, including synthetic data generated in part by the Chinese model Kimi K2.5, a detail that underscores just how intertwined the global AI supply chain has become.

Murati’s team is clearly betting that flexibility will win out over raw scores. They’re previewing a smaller model, Inkling-Small, which actually beats its bigger sibling on a few benchmarks like GPQA Diamond, and they’ve built in a continuously adjustable “thinking effort” knob that lets users trade cost for performance on the fly. It’s a pragmatic, developer-focused pitch from a startup that clearly knows it can’t win a headline war with the Chinese labs on overall intelligence. The question now is whether enterprises will pay a premium for a model that’s savvy but sometimes makes things up. That’s a tough sell, even with Murati’s pedigree behind it.

💡 Key Takeaways

  1. Inkling leads all US open-weights models on the Artificial Analysis Intelligence Index but still trails China's best by a significant margin.
  2. A shockingly high hallucination rate of 63 percent makes the model a risky choice for any application that demands strict factual accuracy.
  3. Thinking Machines is explicitly not competing on raw model intelligence, instead betting its business on fine-tuning and token efficiency as a platform play.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles