AI Pulse by Inblix

GPT-5.6 Luna gets 80% price cut as OpenAI hits 1B users and chases abundance

OpenAI Blog · Jul 31, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: GPT-5.6 Luna gets 80% price cut as OpenAI hits 1B users and chases abundance

OpenAI just made its most aggressive pricing move yet, slashing the cost of GPT-5.6 Luna by 80 percent to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra saw a 20 percent reduction, landing at $2 and $12 respectively. The company also introduced a Fast mode for its top-tier Sol model that delivers up to 2.5x the speed at double the price with no loss in intelligence.

This is not a promotional stunt. It is the latest proof point in what the company calls an “abundance” cycle—a flywheel where cheaper intelligence drives broader adoption, which generates revenue and real-world feedback, which funds the next generation of research and infrastructure. The scale is staggering: OpenAI’s models now reach more than one billion active users and over two million businesses. Six months after signing up, people send roughly 50 percent more messages daily and use ChatGPT for about twice as many types of work.

The economics are shifting from cost-per-token to cost-per-outcome. A more capable model that nails a task in one shot can be cheaper than a bargain model that requires three retries and human oversight. OpenAI is betting that customers will optimize for the “maximum useful intelligence at the right price” rather than chasing the lowest headline rate. The approach is already reshaping internal operations—agentic work through Codex now accounts for 99.8 percent of the company’s weekly output tokens, with the Finance team among the heaviest users.

Engineering gains are compounding fast. GPT-5.6 Sol helped optimize its own production software, cutting end-to-end serving costs by 20 percent and boosting token-generation efficiency by more than 15 percent through improved speculative decoding. More striking still, system-level improvements to reasoning and context management raised Sol’s score on the ARC-AGI-3 benchmark from 13.3 percent to 38.3 percent while using six times fewer output tokens. The model did not change. The scaffolding around it did. That distinction matters: intelligence alone is not enough. The routing, context handling, and product design surrounding the model determine whether that intelligence translates into something useful—and at what cost.

💡 Key Takeaways

  1. System-level improvements to context management nearly tripled GPT-5.6 Sol's ARC-AGI-3 score while slashing output tokens by 6x, proving the wrapper can matter more than the model.
  2. OpenAI's internal agentic work through Codex now represents 99.8% of weekly output tokens, signaling that autonomous task completion has eclipsed simple Q&A inside the company.
  3. A cheaper model that needs repeated attempts can cost more than an expensive one that gets it right the first time—shifting the value metric from price-per-token to cost-per-successful-outcome.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles