GPT-4o can now generate images — and it's a whole different beast
Curated by the Inblix editorial team
OpenAI just tore up the playbook on AI image generation. They’ve embedded a new imaging engine directly into the architecture of GPT-4o, their omnimodal model, and the jump in capability from the old DALL·E 3 series is not incremental — it’s generational. The company published an addendum to the GPT-4o system card detailing what this thing can do, and it reads like a list of things previous models simply couldn’t.
Photorealism is table stakes now. The real shift is in control and comprehension. Because image generation lives deep in the model’s architecture rather than sitting as a separate tool, GPT-4o can use its entire knowledge base when creating visuals. It reliably renders text in images — a persistent headache for earlier systems that often produced garbled alphabet soup. It takes images as inputs and transforms them with a precision that suggests it actually understands what it’s looking at. The system card describes outputs that are not just prettier but genuinely useful, the kind of thing you’d build a workflow around rather than just play with.
None of this happens in a safety vacuum, of course. OpenAI is leaning hard on the infrastructure it built for DALL·E and Sora, applying lessons learned from both deployments. But they’re upfront that new capabilities mean new problems. The addendum focuses on what they call marginal risks — the fresh dangers that emerge specifically because this model is so much more capable. A model that can convincingly alter photographs and embed legible text opens doors that a clumsy image generator never could. The document outlines the mitigations they’ve put in place, though the details of those guardrails will determine whether this feels like a tool or a liability.
What’s genuinely different here is the integration. This isn’t a separate image model bolted onto a chatbot. It’s the same model thinking through images the way it thinks through text, drawing on the same knowledge and following the same nuanced instructions. That architectural choice means the system can be subtle in ways that feel almost expressive — a word that doesn’t normally belong in the same sentence as image generation. Whether that subtlety becomes a superpower or a headache depends on what people do with it, and OpenAI seems acutely aware they’re handing over something sharper than before.
💡 Key Takeaways
- Unlike DALL·E 3, GPT-4o's image generation is embedded natively in the model architecture, letting it leverage the full knowledge base for more coherent and useful outputs.
- The model can now reliably render readable text within images, solving a long-standing failure mode that made previous generators impractical for many professional use cases.
- OpenAI explicitly acknowledges that these new capabilities introduce 'marginal risks' beyond what they faced with DALL·E and Sora, and has built new mitigations specific to photorealistic and text-rendering features.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.