OpenAI's Ultrafast mode hits 750 tokens per second on Cerebras chips
Curated by the Inblix editorial team
OpenAI just answered a complaint that’s been simmering since ChatGPT first went viral: the thing can feel sluggish when you’re waiting on a complex answer. The company’s new Ultrafast mode, announced Thursday, claims to push its latest flagship model, GPT-5.6 Sol, to 14 times standard processing speed — hitting up to 750 output tokens per second. For context, that’s the kind of throughput where a 500-word response materializes in roughly the time it takes to glance at your phone.
The speed isn’t coming from a smaller, dumber model. OpenAI made that point explicitly: “Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.” The distinction matters. Anthropic’s Claude has offered a fast mode for a while, but it doesn’t approach these numbers. OpenAI is claiming you can have the big model and the speed, not a tradeoff between the two.
The catch is the hardware. Ultrafast runs on chips from Cerebras, the wafer-scale chipmaker that’s been quietly positioning itself as the anti-Nvidia for inference workloads. That partnership is the real story here — it signals OpenAI is serious about diversifying its compute stack beyond the GPU supply chain everyone else is fighting over. Cerebras’ architecture is genuinely different: instead of thousands of small chips working in parallel, it uses one massive chip, which eliminates a lot of the communication overhead that slows down traditional clusters.
Access is limited for now — preview only, and OpenAI says it’s restricted to “a small group of customers” with expansion tied to capacity growth. The company is targeting workflows where latency actually costs money: incident response, customer support, financial market analysis, e-commerce. Anyone who’s watched a trading bot hesitate or a support chatbot leave a customer hanging knows exactly why those use cases come first. If the capacity scales and the pricing isn’t absurd, this could reset expectations for what “real-time AI” means across the enterprise.
💡 Key Takeaways
- OpenAI's Ultrafast mode delivers 14x standard processing speed for GPT-5.6 Sol, reaching up to 750 output tokens per second.
- The speed boost comes from Cerebras wafer-scale chips, signaling OpenAI's push to diversify beyond Nvidia-dominated GPU infrastructure.
- Ultrafast targets latency-sensitive enterprise workflows like incident response, financial analysis, and customer support where seconds translate directly to dollars.
- Preview access is limited to a small customer group, with broader rollout contingent on Cerebras capacity expansion.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.