ChatGPT Becomes a Safer Guide Through Difficult Times

ChatGPT Becomes a Safer Guide Through Difficult Times

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
30. 10. 2025
3 minutes reading · 7 views
ChatGPT Becomes a Safer Guide Through Difficult Times

OpenAI has updated ChatGPT's default model to better recognize and support people in moments of distress. They collaborated with more than 170 mental health experts with real-world clinical experience. Thanks to them, the model now more reliably detects signs of distress, de-escalates conversations, and guides people toward real-world support. This reduced responses that did not match the desired behavior by 65–80%. In addition, they expanded access to crisis hotlines, redirected sensitive conversations from other models to safer versions, and added gentle reminders to take breaks during long sessions.

OpenAI believes that ChatGPT can offer support in processing emotions and encourage people to reach out to friends, family, or professionals. The updates focus on three areas: mental health issues such as psychosis or mania, self-harm and suicide, and emotional reliance on AI. In the future, they will also add emotional reliance and non-suicidal mental health distress to the standard evaluations for new models.

Improvement process

These changes build on existing principles in the Model Spec, where OpenAI emphasized supporting real-world relationships, avoiding affirming ungrounded beliefs associated with mental distress, responding safely and empathetically to signs of delusion or mania, and paying attention to indirect signals of self-harm or suicidal thoughts.

They followed five stages to make these improvements: defining the problem by mapping types of potential harm, measuring it using evaluations, data from real-world conversations, and user research, validating the approach with external experts, mitigating risks through model post-training and product interventions, and finally continuing to measure and iterate. They created detailed taxonomies describing the characteristics of sensitive conversations and the model's ideal behavior. The result? A model that responds more reliably to signs of psychosis, mania, suicidal thoughts, self-harm, or unhealthy emotional attachment.

What the measurements revealed

OpenAI defined the areas of concern and quantified their scale. Mental health symptoms are common, but conversations involving risks such as psychosis, mania, or suicidal thoughts are extremely rare—around 0.07% of weekly active users and 0.01% of messages indicate distress associated with psychosis or mania. The new GPT-5 model reduced non-compliant responses by 65% in production. Experts found that GPT-5 reduced undesirable responses by 39% compared with GPT-4o across 677 challenging conversations. In automated evaluations covering more than 1,000 challenging cases, it achieved 92% compliance with the desired behavior, compared with 27% for the previous version of GPT-5.

For self-harm and suicide, they estimate that 0.15% of weekly active users show explicit indicators of suicide planning and that 0.05% of messages contain explicit or implicit signals. The new model reduced non-compliant responses by 65%, while experts observed a 52% decrease compared with GPT-4o across 630 conversations. Automated evaluations show 91% compliance, compared with 77% for the previous GPT-5. In long conversations, it maintains reliability above 95%.

For emotional reliance, they estimate that 0.15% of weekly users and 0.03% of messages show elevated levels of attachment. The updates reduced non-compliant responses by 80%, while experts recorded a 42% decrease compared with GPT-4o across 507 conversations. Automated evaluations show 97% compliance, compared with 50% for the previous version.

Fewer non-compliant responses

Collaboration with experts

OpenAI created the Global Physician Network, comprising nearly 300 physicians and psychologists from 60 countries. More than 170 of them helped write ideal responses, analyze model responses, evaluate safety, and provide feedback. Psychiatrists and psychologists reviewed more than 1,800 responses in serious situations and found a 39–52% decrease in undesirable responses compared with GPT-4o. Inter-rater agreement among experts was 71–77%. They also collaborated on evaluations such as HealthBench for internal testing.

OpenAI plans to continue developing taxonomies and measurement systems, with details provided in the addendum to the GPT-5 system card.

Experts' assessment of inappropriate responses

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok