AI Pulse by Inblix

OpenAI drops o3 and o4-mini, models that think with tools

OpenAI Blog · Jul 14, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI drops o3 and o4-mini, models that think with tools

OpenAI just launched o3 and o4-mini, and the headline feature isn’t just better reasoning. These models use tools directly inside their chain of thought. That means during the messy, internal process of figuring out an answer, the model can pause, crop an image, run a Python script, search the web, or pull from memory. It’s not just thinking out loud—it’s reaching for a calculator mid-sentence.

This is a meaningful shift in architecture, not just a benchmark bump. The models are trained with large-scale reinforcement learning on those chains of thought, which OpenAI says opens new avenues for safety. They’re calling it “deliberative alignment”—the ability for the model to reason about the company’s safety policies in context when it hits a potentially unsafe prompt, rather than relying solely on pre-baked refusals.

On the safety front, the launch is notable for being the first release under Version 2 of OpenAI’s Preparedness Framework. The Safety Advisory Group reviewed the evaluations and determined neither o3 nor o4-mini hits the “High” risk threshold in biological, chemical, cybersecurity, or AI self-improvement categories. The company also released two system card addendums: one covering Codex and another for an o3 Operator. That’s a level of documentation that suggests they’re bracing for regulatory and public scrutiny.

There’s always a gap between a system card and the real world, and I’ll be watching to see how these models handle edge cases when tool use goes sideways. But the core idea is genuinely practical: a model that can look something up or write a quick script to check its own work, all before it ever gives you an answer. That changes what you can trust it with.

💡 Key Takeaways

  1. OpenAI's o3 and o4-mini integrate tools like Python, web search, and image analysis directly into their internal reasoning process, not just as post-hoc add-ons.
  2. The models use 'deliberative alignment' to reason about OpenAI's safety policies in context when handling potentially harmful prompts.
  3. This is the first launch evaluated under OpenAI's stricter Preparedness Framework v2, and neither model crossed the 'High' risk threshold in any tracked danger category.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles