OpenAI Releases 'MentalHealthBench' with Over 80 Experts from 22 Countries
Key point
OpenAI has released 'MentalHealthBench', an open benchmark developed in collaboration with over 80 experts from 22 countries to evaluate AI's ability to handle mental health conversations.
Details
OpenAI has released MentalHealthBench, an open benchmark designed to evaluate AI systems' ability to handle mental health conversations. Unlike previous evaluations that primarily focused on emergencies, this benchmark covers the entire spectrum of mental health, from everyday stress to acute crises, measuring model safety, context awareness, and respect for user autonomy.
Expert-Based Evaluation Framework
This benchmark was developed in collaboration with over 80 licensed psychologists and psychiatrists from 22 countries. Experts reviewed synthetic conversation scenarios and created weighted rubrics to evaluate specific aspects of model responses. The criteria have weights ranging from -10 to +10 points, assigning positive scores to beneficial behaviors and negative scores to harmful behaviors. The final evaluation uses GPT-5.6 Sol as an automated judge to score model responses according to the expert-created criteria.
Diverse Scenarios and User Types
MentalHealthBench includes various user types such as adults, adolescents aged 13–17, caregivers, and clinicians, covering three severity levels: non-acute (everyday emotional conversations), advanced (severe distress), and emergency (immediate safety concerns). It also reflects 19 languages and approximately 20 mental health sub-specialties to conduct evaluations that consider cultural context and diversity.
Comparative Analysis with User Perspectives
To analyze the difference between expert guidelines and the usefulness perceived by actual users, OpenAI conducted a separate survey of 44 adults from 16 countries. The analysis revealed that users prioritize 'practical next steps' and 'tone' more than experts, while experts place greater emphasis on 'context gathering' and 'interpreting ambiguous situations'. This comparison provides insights necessary for AI support to balance clinical accuracy and user-friendliness.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.