OpenAI tightens safety rules, admits rivals could force its hand
Curated by the Inblix editorial team
OpenAI just published a significant update to its Preparedness Framework, the internal rulebook that governs how the company tracks and mitigates catastrophic risks from its most advanced models. The refresh isn’t just bureaucratic housekeeping; it signals a sharper, more operationalized approach to safety at a time when the frontier is accelerating. The company is drawing a brighter line between capabilities it actively tracks with mature evaluations — biology, cybersecurity, and AI self-improvement — and a new set of ‘Research Categories’ where the threat models aren’t baked yet. These emerging buckets include long-range autonomy, the ability for AI to undermine its own safeguards, and nuclear and radiological risks. It’s a frank admission that they don’t have a handle on everything, and they’re scrambling to build the measurement tools for risks they see coming around the corner.
The framework now hinges on two clear thresholds: ‘High’ capability, which amplifies existing dangers, and ‘Critical’ capability, which could introduce entirely novel paths to severe harm. Hitting the ‘High’ bar means safeguards must be in place before a model ships; hitting ‘Critical’ means those protections are required even during internal development. A cross-functional Safety Advisory Group reviews the evidence and makes recommendations to leadership, who have the final say. The company is also trying to keep pace with the speed of model updates, leaning more on automated evaluations to match the cadence of post-training improvements that don’t require a full retraining run. Expert-led deep dives will continue, but the shift acknowledges that manual testing alone can’t keep up.
Persuasion risks are notably getting booted out of this framework entirely. OpenAI says those will be handled by other means — its Model Spec, bans on political campaigning, and ongoing investigations into covert influence operations. It’s a practical carve-out that separates the messy, context-dependent world of information manipulation from the more concrete, dual-use dangers of bio and cyber capabilities. This move likely reflects both the difficulty of measuring persuasion objectively and a desire to keep the Preparedness Framework focused on truly existential, rather than societal, harms.
Perhaps the most eyebrow-raising detail is a clause addressing third-party behavior. If another frontier lab releases a high-risk system without comparable guardrails, OpenAI reserves the right to reconsider its own restrictions. The language is careful — they’d first need to confirm the threat landscape has actually changed — but the implication is clear. In an all-out race, a unilateral safety tax becomes harder to justify. It’s a pragmatic, slightly unnerving admission that even the most well-intentioned safety framework exists within a competitive ecosystem that doesn’t always reward restraint.
💡 Key Takeaways
- OpenAI now splits risk tracking into mature 'Tracked Categories' like bio and cyber, and newer 'Research Categories' including long-range autonomy and AI undermining safeguards, where threat models are still under construction.
- A new two-tier system ('High' and 'Critical' capability) dictates exactly when safeguards must be active — before deployment for the former, and during development for the latter — with a Safety Advisory Group making recommendations to leadership.
- Persuasion and influence risks have been explicitly removed from this framework and will be governed by separate policies, including the Model Spec and bans on political campaigning.
- OpenAI explicitly leaves the door open to lowering its own safety requirements if a rival ships a high-risk model without comparable guardrails, after confirming the risk landscape has shifted.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.