OpenAI’s o1 can now “think” its way past safety rules
Curated by the Inblix editorial team
OpenAI has published the system card for its o1 model family, and the headline isn’t just that the thing can code. It’s that the company taught a large language model to reason through its safety policies in real time. That approach—what OpenAI calls “deliberative alignment”—lets o1 models chew on a user’s prompt with a long chain-of-thought before answering, in theory helping them spot jailbreak attempts and refuse harmful requests more gracefully than previous models. The company claims state-of-the-art scores on internal benchmarks for resisting illicit advice, avoiding stereotyped responses, and shutting down known jailbreaks. That’s a genuine shift: instead of relying entirely on post-training guardrails, the model is actively interpreting the rules in context.
But the safety card isn’t all green. OpenAI’s Preparedness Framework scores o1 as “Medium” on both chemical, biological, radiological, and nuclear (CBRN) risks and persuasion. That’s a notch above the “Low” ratings for cybersecurity and model autonomy, but it’s still the kind of label that makes policy folks nervous. The report doesn’t go into granular detail about what a “Medium” persuasion risk looks like in practice, but it’s hard not to read between the lines: a model that can reason better can also argue better. OpenAI itself notes that heightened intelligence cuts both ways, which is a refreshingly blunt admission from a company that usually leads with the utopian pitch.
The training data recipe is a mix of public web data, open-source datasets, and proprietary content from paywalled partnerships. That last bucket is worth watching—it’s a nod to the quiet content deals that have become table stakes for frontier labs, and it raises the inevitable question of whose copyrighted work is being digested without explicit consent. OpenAI says it uses rigorous filtering and safety classifiers to strip personal information and block harmful content, including CSAM. But filtering claims are cheap; what matters is whether the pipeline actually catches edge cases before they reach users.
What’s genuinely different here is the evaluation timeline. The report includes numbers from two checkpoints: a near-final version that went through external red teaming and the December 5 release that added format and instruction-following tweaks on top of the same base model. That’s a level of transparency that safety researchers have been begging for, and it makes the claims more testable. Still, I’ll believe the robustness when I see it survive a few months of creative users poking at it on social media. A model that reasons about safety is a model that can, given enough bad-faith prodding, reason its way around it.
💡 Key Takeaways
- o1 uses “deliberative alignment” to reason about safety policies before responding, achieving top scores on jailbreak and harmful content benchmarks.
- The model scores “Medium” on CBRN and persuasion risks, signaling that stronger reasoning capability introduces new kinds of misuse potential.
- Training data includes proprietary, paywalled content from partnerships, raising transparency and copyright questions despite OpenAI’s filtering claims.
- OpenAI tested two distinct checkpoints, sharing data from both external red teaming and final release evaluations—a notable step toward verifiable safety claims.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.