AI Pulse by Inblix

OpenAI drops o3-mini: a faster, smarter STEM specialist

OpenAI Blog · Jul 15, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI drops o3-mini: a faster, smarter STEM specialist

OpenAI just rolled out o3-mini, and it’s not just another model update — it’s a clear signal that the company is serious about making reasoning models both faster and cheaper for developers. Previewed in December, this small but mighty model is now live in ChatGPT and the API, and it’s specifically tuned to excel at science, math, and coding tasks. What makes this release genuinely interesting is the developer-centric flexibility: for the first time in a small reasoning model, you get function calling, Structured Outputs, and developer messages, which means it’s actually production-ready from day one. Developers can also dial in the model’s “reasoning effort” — low, medium, or high — letting them trade speed for deeper thought depending on the task. It’s a pragmatic move that acknowledges not every problem needs the AI to sit and ponder for ten seconds.

On the performance front, the numbers are worth paying attention to. OpenAI claims o3-mini with medium reasoning effort matches the much larger o1 model on tough benchmarks like AIME and GPQA, while delivering responses 24% faster than o1-mini. That’s a 7.7-second average response time versus 10.16 seconds. Expert testers preferred o3-mini’s answers over o1-mini’s 56% of the time and observed a 39% drop in major errors on difficult real-world questions. It’s not a revolution; it’s a refinement. But the efficiency gains are real, and they build on OpenAI’s broader push to slash costs — the company notes per-token pricing has fallen 95% since GPT-4 launched. That’s the kind of stat that makes CFOs, not just ML engineers, sit up.

Accessibility is the other headline here. For the first time, free-tier ChatGPT users can try a reasoning model by selecting “Reason” in the composer. Paid users get a triple rate limit boost — from 50 messages a day with o1-mini to 150 with o3-mini — and Pro subscribers get unlimited access to both the standard and high-intelligence variants. There’s even a search integration prototype baked in, letting the model pull up-to-date web sources, though it’s clearly early days for that feature. One notable gap: o3-mini can’t handle vision tasks, so if you’re working with images, you’re still stuck with o1.

I’m less interested in the safety boilerplate — yes, they used deliberative alignment and red-teaming, and yes, it outperforms GPT-4o on jailbreak evals — and more in what this says about OpenAI’s strategy. They’re not chasing the single biggest model anymore. They’re building a family of specialized tools at different price points, and o3-mini is the sharp, fast, affordable option for coders and quants. The question now is whether developers actually adopt these reasoning effort sliders or just crank everything to high and complain about latency. My bet’s on the latter, at least until the sticker shock kicks in.

💡 Key Takeaways

  1. o3-mini matches o1's performance on STEM benchmarks like AIME and GPQA while running 24% faster than o1-mini, with an average 7.7-second response time.
  2. The model introduces a three-tier reasoning effort slider (low, medium, high), giving developers granular control over the speed-versus-accuracy tradeoff for the first time in a small reasoning model.
  3. Free ChatGPT users can now access a reasoning model — a first for OpenAI — while paid users see rate limits triple to 150 messages per day.
  4. OpenAI claims a 95% reduction in per-token pricing since GPT-4, positioning o3-mini as the next step in making high-end reasoning economically viable at scale.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles