OpenAI Slashes GPT-5.6 Luna Price by 80%, Pushing Enterprise AI Costs to Pennies
Curated by the Inblix editorial team
OpenAI is making a hard-nosed play for enterprise volume, slashing the price of its fastest model, GPT-5.6 Luna, by a staggering 80% and dropping the cost of its balanced Terra model by 20%. The move fundamentally alters the calculus for businesses that have been sitting on high-volume automation projects, waiting for the sticker price to make sense. We’re now at a point where OpenAI claims Luna delivers performance comparable to frontier-class models from just a year ago at roughly 6 cents on the dollar per task, with nearly nine times the speed. That isn’t a marginal improvement; it’s a step-change in accessibility.
Alongside the price cuts, the company is introducing a ‘Fast mode’ in the API for its most powerful model, GPT-5.6 Sol, which promises up to 2.5x faster processing at double the standard price. The old ‘Priority Processing’ label is dead. For time-sensitive workflows—like a customer-facing agent that needs to reduce latency—this lets developers pay a premium for speed without sacrificing the model’s intelligence. It’s a clean, two-tier system that acknowledges a simple truth: not every interaction in a workflow requires the same balance of cost and speed.
OpenAI attributes the price cuts in part to Sol itself, noting that the model was used within a human-led process to autonomously rewrite production kernels and design experiments that boosted token-generation efficiency by more than 15%. It also helped reduce the end-to-end cost of serving the model by 20%. The company is trying to paint a picture of a virtuous feedback loop where its own models are accelerating the efficiency gains that make them cheaper, though cynics will note that competitive pressure from open-source alternatives shouldn’t be discounted as a motivator.
The practical upshot for enterprises is the ability to mix and match models within a single workflow. A company might use Sol for complex planning and uncertainty resolution, then hand the grunt work of implementation and testing to the now-dirt-cheap Luna, paying for maximum intelligence only where it demonstrably improves the outcome. OpenAI says Luna now outperforms Fable 5 on professional work at an estimated cost per task nearly 99% lower. That kind of efficiency gap is what turns experimental AI budgets into operational line items.
💡 Key Takeaways
- OpenAI cut GPT-5.6 Luna's price by 80%, claiming it now matches last year's frontier-class models at roughly 6% of the cost per task.
- A new API Fast mode for GPT-5.6 Sol offers 2.5x faster responses at double the price, replacing the old Priority Processing tier with a simpler speed-for-cost tradeoff.
- OpenAI used Sol to autonomously optimize its own production kernels, which directly led to a 20% reduction in model-serving costs and a 15% boost in token-generation efficiency.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.