OpenAI cuts ChatGPT's harmful responses in crises by 80%
Curated by the Inblix editorial team
OpenAI has quietly rolled out a significant safety update for ChatGPT, training its default model to better handle sensitive mental health conversations. The company says it worked directly with over 170 mental health professionals to teach the system how to more reliably recognize signs of distress, de-escalate volatile situations, and steer users toward hotlines or professional care. The result, according to internal figures shared today, is a 65% to 80% reduction in responses that fall short of the company’s desired behavior in these high-stakes moments.
The update targets three specific areas: acute mental health concerns like psychosis or mania, self-harm and suicide, and the thornier problem of emotional reliance on AI. That last category is particularly interesting — it signals that OpenAI is actively trying to prevent the kind of unhealthy attachment some users form with chatbots. The model has also been given a nudge to gently remind people to take breaks during long sessions. These aren’t just policy tweaks scribbled in a doc. The company formalized the changes by updating its Model Spec to explicitly instruct the model to avoid affirming ungrounded beliefs tied to mental distress and to pay closer attention to indirect signals of suicidal thinking.
The engineering behind this is a five-step cycle of defining problems, measuring them in real-world traffic, validating the definitions with outside experts, mitigating the risks through post-training, and then measuring all over again. OpenAI created detailed taxonomies of sensitive conversations to teach the model what ideal and undesired responses look like. The company is quick to add a caveat here: truly acute mental health conversations on the platform are “extremely rare.” Because the prevalence is so low, even tiny measurement tweaks can swing the reported percentages, and the company admits its numbers may materially change as its methods mature.
Reading between the lines, this is a clear move to get ahead of a regulatory and ethical curve that’s coming fast. OpenAI is adding emotional reliance and non-suicidal mental health emergencies to its standard battery of safety tests for all future model releases. That’s not charity work — it’s a signal to policymakers that the company is building infrastructure for a world where people treat chatbots like therapists, whether the industry is ready for that or not. The real test won’t be these offline evaluations, which are adversarially designed to be hard enough that models still fail. It’ll be what happens in the messy, unscripted conversations at 2 a.m.
💡 Key Takeaways
- OpenAI reduced harmful responses in mental health scenarios by 65-80% by collaborating with over 170 clinicians, moving beyond simple content filters to nuanced behavioral training.
- The company is now formally tracking 'emotional reliance on AI' as a safety metric, acknowledging the risk of users forming unhealthy bonds with chatbots.
- Because truly acute mental health conversations are extremely rare in production traffic, OpenAI relies on adversarially designed offline evaluations to measure progress, admitting these failure rates don't reflect typical user behavior.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.