ChatGPT Strengthens Awareness of Sensitive Conversation Context
Key point
OpenAI has strengthened ChatGPT's awareness of sensitive conversation context and its safe responses.
Details
OpenAI has improved ChatGPT's model policy, training, evaluation, and monitoring systems so it can better catch risk signals that emerge gradually within and across conversations. Rather than looking only at a single message, it now considers prior context together, and when signs of suicide/self-harm or harm to others appear, it treats these as rare but cases requiring extra caution, responding with more careful refusals or guiding users toward safer alternatives.
This work was carried out based on more than 2 years of collaboration with mental health and safety experts. Psychiatrists and psychologists from the Global Physicians Network helped refine when to create safety summaries, how much prior context to reflect, and how long that context should be retained.
Safety summaries, produced by a model trained for safety reasoning, retain only brief, factual records of past safety-related facts, are stored for a limited period, and are not used for general personalization or long-term memory purposes. This change builds on the safe completion approach, which refuses only the unsafe portion of a request while answering as safely as possible within the allowed scope.
In internal evaluations, performance improved significantly in scenarios where risk emerges late.
- In long single conversations, safe responses improved by 50% for suicide/self-harm and 16% for harm to others.
- On GPT-5.5 Instant, ChatGPT's current default model, improvements were 52% for harm to others and 39% for suicide/self-harm.
- Across more than 4,000 evaluations, safety summaries scored 4.93/5 for relevance and 4.34/5 for factual accuracy.
- In everyday conversations, there was no drop in response quality or any meaningful difference in preference.
OpenAI stated that this approach could be extended to other high-risk areas such as biology or cyber safety in the future.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.