Our Commitment to Community Safety
Key point
OpenAI announced it is strengthening ChatGPT's response to violence and self-harm risks along with its enforcement system.
Details
OpenAI stated that ChatGPT maintains a balance of refusing requests that could lead to violence or self-harm, while allowing neutral questions for news, history, education, and prevention purposes.
- Based on the Model Spec, actionable violent instructions, tactics, and plans are blocked.
- The detection system has been strengthened to better capture subtle risk signals by analyzing long conversations and patterns across conversations.
- The company continues to refine the balance of safety, privacy, and accessibility with advice from psychologists, psychiatrists, civil liberties experts, and law enforcement experts.
In crisis situations, the system guides users to region-specific crisis resources, mental health professionals, and connections with trusted acquaintances, and recommends emergency help in serious cases.
On the operational side, an automatic detection system uses classifiers, reasoning models, hash matching, blocklists, and more to filter suspicious cases. Afterward, human review verifies context under restricted access privileges and security/confidentiality rules.
When a violation is confirmed, the company applies its Usage Policies' zero-tolerance principle, which can include account deactivation, blocking associated accounts, and blocking the creation of new accounts. Users can appeal enforcement actions, and the company reviews them again. If an imminent and credible risk of violence is identified, law enforcement is notified.
Parental Controls, introduced last fall, allow parents to link teen accounts to adjust safety settings, without being able to view conversation content directly. When the system and human review detect acute crisis signals, notifications are sent via email, SMS, and push notifications, and a trusted contact feature designated by adult users is also planned to be introduced soon.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.