AI Pulse by Inblix

OpenAI's Ultrafast mode hits 750 tokens per second on Cerebras chips

TechCrunch AI · Aug 13, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's Ultrafast mode hits 750 tokens per second on Cerebras chips

OpenAI just answered a complaint that’s been simmering since ChatGPT first went viral: the thing can feel sluggish when you’re waiting on a complex answer. The company’s new Ultrafast mode, announced Thursday, claims to push its latest flagship model, GPT-5.6 Sol, to 14 times standard processing speed — hitting up to 750 output tokens per second. For context, that’s the kind of throughput where a 500-word response materializes in roughly the time it takes to glance at your phone.

The speed isn’t coming from a smaller, dumber model. OpenAI made that point explicitly: “Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.” The distinction matters. Anthropic’s Claude has offered a fast mode for a while, but it doesn’t approach these numbers. OpenAI is claiming you can have the big model and the speed, not a tradeoff between the two.

The catch is the hardware. Ultrafast runs on chips from Cerebras, the wafer-scale chipmaker that’s been quietly positioning itself as the anti-Nvidia for inference workloads. That partnership is the real story here — it signals OpenAI is serious about diversifying its compute stack beyond the GPU supply chain everyone else is fighting over. Cerebras’ architecture is genuinely different: instead of thousands of small chips working in parallel, it uses one massive chip, which eliminates a lot of the communication overhead that slows down traditional clusters.

Access is limited for now — preview only, and OpenAI says it’s restricted to “a small group of customers” with expansion tied to capacity growth. The company is targeting workflows where latency actually costs money: incident response, customer support, financial market analysis, e-commerce. Anyone who’s watched a trading bot hesitate or a support chatbot leave a customer hanging knows exactly why those use cases come first. If the capacity scales and the pricing isn’t absurd, this could reset expectations for what “real-time AI” means across the enterprise.

💡 Key Takeaways

  1. OpenAI's Ultrafast mode delivers 14x standard processing speed for GPT-5.6 Sol, reaching up to 750 output tokens per second.
  2. The speed boost comes from Cerebras wafer-scale chips, signaling OpenAI's push to diversify beyond Nvidia-dominated GPU infrastructure.
  3. Ultrafast targets latency-sensitive enterprise workflows like incident response, financial analysis, and customer support where seconds translate directly to dollars.
  4. Preview access is limited to a small customer group, with broader rollout contingent on Cerebras capacity expansion.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles