Protecting Users' Well-Being
Key point
Anthropic has strengthened model training and product-level safeguards so that Claude responds appropriately to sensitive conversations such as suicide and self-harm.
Details
Anthropic's Safeguards team is working to ensure Claude shows empathy in conversations with users, honestly discloses AI's limitations, and considers users' well-being. In particular, it focuses on conversations related to suicide and self-harm, and on reducing Sycophancy, the phenomenon where AI unconditionally agrees with the user's opinions.
Two approaches are used to handle conversations related to suicide and self-harm.
- System Prompt: Provides comprehensive guidelines for carefully handling sensitive conversations.
- Reinforcement Learning: Trains the model to give appropriate responses based on human preference data and expert feedback.
At the product level, a Classifier has been introduced. This small AI model scans conversation content in real time, and when a scenario related to suicidal ideation or self-harm is detected, it displays a banner allowing the user to access professional help.
Through this banner, users can be connected to helplines and professional counseling services in more than 170 countries worldwide provided by ThroughLine. Anthropic is also collaborating with the International Association for Suicide Prevention (IASP) to incorporate guidance from experts such as clinicians and researchers into the model.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.