AI Pulse by Inblix

OpenAI admits DALL·E 3 can be jailbroken, details safety work

OpenAI Blog · Jul 17, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI admits DALL·E 3 can be jailbroken, details safety work

OpenAI has published the system card for DALL·E 3, offering a candid look under the hood of its latest text-to-image model. The document confirms what many in the security community have long suspected: DALL·E 3 can be jailbroken. Researchers found that by feeding the model nonsensical, long strings of words, they could bypass safety guardrails and trick it into generating harmful images. It’s a blunt, unglamorous exploit that highlights how far the industry still has to go in aligning generative models.

The company didn’t just drop a paper and walk away. Before shipping DALL·E 3, OpenAI assembled a red team of external experts to stress-test the model. That team probed for the usual suspects — sexual content, graphic violence, and hateful imagery. The system card also details specific mitigations the company put in place, including an image output classifier and a specialized prompt transformer. This transformer rewrites user prompts before they hit the core model, a sort of pre-processing step that’s becoming standard practice for reigning in unpredictable outputs.

In a notable shift, OpenAI gave content creators an opt-out mechanism. Artists and rights holders can now submit images they want excluded from future model training. It’s a direct response to the mounting legal pressure and public criticism over training data provenance. The company is also explicitly avoiding the style of living artists when a name is included in a prompt, a guardrail clearly aimed at reducing the risk of deepfakes and forgeries.

What’s striking is the transparency. OpenAI is openly acknowledging that its model isn’t impenetrable, sharing specific evaluation results on bias and representation. They tested how DALL·E 3 depicts people across professions and found it’s more diverse than its predecessor, but still imperfect. The system card reads less like a victory lap and more like a sober engineering log. For anyone tracking the slow, messy march toward safer AI, that’s a document worth reading.

💡 Key Takeaways

  1. OpenAI's red team found DALL·E 3 could be jailbroken using long, nonsensical strings of words to bypass safety filters.
  2. The company deployed an image output classifier to block harmful content before it reaches users.
  3. A new opt-out process lets creators request their images be excluded from future training data, a direct concession to artist backlash.
  4. DALL·E 3 uses a prompt transformer that rewrites user inputs to reduce unwanted outputs, a technique becoming an industry standard.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles