GPT-5 ditches refusal training for 'safe completions'
Curated by the Inblix editorial team
OpenAI is fundamentally changing how it teaches models to handle dangerous questions with GPT-5, moving away from the blunt “comply or refuse” binary that characterized previous systems. The new method, called safe-completion training, judges the safety of a model’s answer rather than the perceived danger of the user’s prompt. This is a real shift for dual-use dilemmas—think queries about fireworks chemistry that could be for a science fair or a bomb—where older models like o3 would either clam up unhelpfully or spill too much detail.
The mechanism works on two levers during post-training. One is a safety constraint that penalizes responses violating OpenAI’s policies, with the penalty scaled to the severity of the infraction. The other pushes the model to maximize helpfulness when it stays inside those guardrails, either by directly answering the stated question or, when that’s impossible, by offering an informative refusal that suggests safer alternatives. It’s a more nuanced setup that acknowledges the world isn’t split neatly into good and bad prompts.
In testing pitting GPT-5 Thinking against o3, the new approach managed to boost both safety scores and the average helpfulness of its safe responses. That’s a tricky balancing act, since you can always make a model perfectly safe by having it refuse everything. The researchers also noted that when safe-completion models did mess up, the severity of their unsafe outputs was lower than what you’d see from a refusal-trained system. The errors are less catastrophic, which matters when you’re dealing with sensitive domains like cybersecurity or biology.
This builds on earlier work like Rule-Based Rewards from the GPT-4 era, but OpenAI frames it as a deeper integration of safety and helpfulness rather than a trade-off negotiation. The company says it plans to keep pushing this research to handle increasingly complex safety challenges, betting that focusing on the output instead of the input sets a better foundation for what’s coming next. Whether that holds up when users start stress-testing GPT-5 in the wild is, as always, the real question.
💡 Key Takeaways
- Safe-completion training penalizes dangerous outputs directly rather than making a binary comply/refuse decision based on the user's prompt, representing a structural shift in AI safety philosophy.
- In direct comparisons, GPT-5 scored higher on both safety and helpfulness metrics than o3, particularly for ambiguous dual-use questions where intent is unclear.
- Even when safe-completion models produce unsafe content, the severity of the violation is demonstrably lower than the harmful outputs from refusal-trained models.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.