AI Pulse by Inblix

OpenAI's GPT-4o mini is 60% cheaper than GPT-3.5 Turbo

OpenAI Blog · Jul 16, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's GPT-4o mini is 60% cheaper than GPT-3.5 Turbo

OpenAI has rolled out GPT-4o mini, and calling it a “small” model almost undersells what it can do. This is a deliberate shot across the bow at competitors like Google’s Gemini Flash and Anthropic’s Claude Haiku, and the numbers back it up. We’re talking 82% on MMLU, a textual reasoning benchmark where Gemini Flash managed 77.9% and Claude Haiku trailed at 73.8%. For math reasoning on MGSM, it hit 87% versus 75.5% for Flash and 71.7% for Haiku. The coding benchmark HumanEval shows a similar spread: 87.2% for GPT-4o mini, a figure that leaves Claude Haiku’s 75.9% and Gemini Flash’s 71.5% in the dust.

The real headline, however, might be the economics. At 15 cents per million input tokens and 60 cents per million output tokens, this model is more than 60% cheaper than the aging GPT-3.5 Turbo. It is literally an order of magnitude more affordable than previous frontier models. That price drop isn’t just a nice-to-have; it fundamentally changes the math on what’s feasible to build. Developers can now chain multiple API calls, feed massive codebases or conversation histories into a 128K context window, and run real-time customer service chatbots without the cost being a dealbreaker.

OpenAI is positioning this as an immediate tool and a future platform. Right now, the API supports text and vision. Audio and video capabilities are on the roadmap. The company also brought in early partners like Ramp and Superhuman, who reported that GPT-4o mini ran circles around GPT-3.5 Turbo for practical tasks like pulling structured data from a crumpled receipt or drafting an email that actually sounds like you from a messy thread history. That kind of real-world validation from enterprise users is a stronger signal than any benchmark.

Safety and security are getting a fresh approach here too. GPT-4o mini is the first model to implement OpenAI’s new “instruction hierarchy” technique, a method designed to make the model more resistant to jailbreaks, prompt injections, and system prompt extractions. That’s a direct response to the headaches developers face when scaling applications. With more than 70 external experts having poked at the model for risks, OpenAI is betting that making a model both cheaper and harder to trick is a combination that will win over the builders who are currently stitching together their own solutions from less reliable parts.

💡 Key Takeaways

  1. GPT-4o mini's 60% price drop compared to GPT-3.5 Turbo makes it economically viable to build applications that chain multiple model calls or process huge context windows.
  2. The model matches or beats named competitors—Gemini Flash and Claude Haiku—on key benchmarks including MMLU, MGSM, and HumanEval.
  3. OpenAI is applying a new 'instruction hierarchy' safety technique to this model first, aiming to directly combat prompt injections and jailbreaks at scale.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles