OpenAI's Ultrafast tier hits 750 tokens/second with GPT-5.6 Sol on Cerebras
Curated by the Inblix editorial team
OpenAI is previewing a new service tier called Ultrafast that runs its GPT-5.6 Sol model up to 14 times faster than standard processing, powered by Cerebras hardware. The headline number: up to 750 output tokens per second. That’s not a small optimization — it’s an order-of-magnitude shift that puts frontier-tier intelligence into latency budgets previously reserved for tiny, specialized models.
The real story here isn’t raw speed. It’s what happens when you stop forcing teams to choose between a smart model and a fast one. OpenAI frames the early use cases around time-sensitive work: engineers triaging outages while logs are still streaming, financial analysts flagging suspicious transactions before conditions shift, and customer support agents resolving multi-step issues mid-conversation without awkward pauses. Even internal OpenAI teams are using it to tighten overnight research batches into interactive, same-day iteration loops.
What’s interesting is the strategic signal. OpenAI has spent years pushing capability ceilings — bigger context windows, better reasoning, more agentic behavior. Ultrafast pushes in a different direction entirely: throughput as a product differentiator. The company is betting that latency, not just intelligence, will determine where AI actually gets deployed in production. If a model can keep pace with a human operator, the argument goes, entirely new workflow categories open up. That’s the pitch behind the commerce example too — answering product questions and resolving checkout issues before a hesitant shopper abandons the cart.
Cerebras deserves real credit here. The wafer-scale chipmaker has been chasing ultra-low-latency inference for years, and this marks its most significant deployment to date: running OpenAI’s most intelligent model, not a distilled or quantized variant. The partnership suggests OpenAI sees value in hardware diversity rather than betting everything on its own inference stack.
Access is limited for now. OpenAI is running a preview with select customers across coding, commerce, financial research, and support, with a waitlist for broader rollout. The company says it will use findings from these production environments to guide deployment as capacity scales. Whether Ultrafast becomes a premium add-on, a default option, or something bundled into existing tiers remains an open question — but the direction is clear. Speed is no longer a consolation prize for using a dumber model.
💡 Key Takeaways
- Ultrafast runs GPT-5.6 Sol at up to 750 output tokens per second, a 14x speedup over standard processing, via Cerebras hardware
- The tier targets time-critical workflows like incident response, fraud detection, and live customer support where latency directly impacts outcomes
- OpenAI is using the limited preview to study how order-of-magnitude speed changes reshape workflows before expanding capacity
- Cerebras is now powering OpenAI's most intelligent model, signaling a meaningful hardware partnership beyond OpenAI's internal inference stack
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.