OpenAI's o3 and o4-mini can now 'think' with images, not just see them
Curated by the Inblix editorial team
OpenAI just dropped a pair of new models—o3 and o4-mini—and for the first time, the chain-of-thought process includes actual visual reasoning. We’re not talking about simply identifying what’s in a picture. These models can actively manipulate images during their internal deliberation, cropping, zooming, rotating, and applying other processing techniques natively, without outsourcing the work to a separate specialized model. It’s a fundamental shift in how the model arrives at an answer, blending visual and textual clues into a single reasoning stream. You can throw a messy photo at it—text upside down, multiple problems on a whiteboard, a poorly framed screenshot of a build error—and the model will reportedly fiddle with the image itself to extract what it needs.
The performance gains are stark, particularly on perception benchmarks that previously stumped multimodal systems. On V*, a visual search benchmark, the new approach hits 95.7% accuracy. The company says that’s largely solving it. They also set new state-of-the-art marks on STEM question-answering tests like MMMU and MathVista, chart reading tasks like CharXiv, and even the cheekily named ‘VLMs are Blind’ benchmark that exposes basic perception failures. This isn’t just a marginal bump; letting the model ‘think with images’ without leaning on web browsing delivered significant leaps across the board, according to their internal evaluations. The models also integrate with tools like Python analysis and web search, creating what OpenAI calls its first ‘multimodal agentic experience.’
Of course, no model launch comes without caveats. OpenAI is refreshingly candid about the current limitations. The reasoning chains can get excessively long, with the model sometimes making redundant or unnecessary tool calls that bloat the thinking process. Basic perception errors still creep in, where a correct manipulation step gets derailed by a simple visual misinterpretation. Reliability across multiple attempts on the same problem can also waver—the model might try one reasoning path that works and another that doesn’t. It’s powerful, but it’s not yet a precision instrument. The company says it’s actively working to make the reasoning more concise and dependable.
What’s genuinely interesting here is the axis of scaling this unlocks. We’ve watched test-time compute expand through longer text-based reasoning chains. Now OpenAI is adding a visual dimension to that compute budget, letting the model spend its inference time not just thinking longer, but looking closer. For anyone who regularly deals with diagrams, handwritten notes, or messy data, the promise of a model that can squint and rotate its way to an answer feels like a practical step forward, not just another benchmark flex. The real test will be whether this ‘thinking with images’ proves robust outside the lab, where the photos are even worse and the stakes are higher.
💡 Key Takeaways
- OpenAI's o3 and o4-mini models can natively crop, zoom, and rotate images during their internal chain of thought without relying on separate specialized models.
- The 'thinking with images' approach achieved 95.7% accuracy on the V* visual search benchmark, which the company says largely solves the test.
- OpenAI acknowledges current flaws including excessively long reasoning chains and basic perception errors that can still derail a correct final answer.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.