OpenAI Releases MentalHealthBench for AI Evaluation
OpenAI's announcement introduces MentalHealthBench, an open benchmark designed to measure how AI systems respond in realistic mental health conversations. The project was co-created with a cohort of more than 80 licensed mental health experts from 22 countries. These experts represented 19 languages and nearly 20 mental health subspecialties.
The benchmark assesses model capabilities across behaviors such as safety, seeking context, preserving user agency, and providing actionable guidance. To create the evaluation, OpenAI used privacy-preserving techniques to generate synthetic mental health conversations involving adults, teens, caregivers, and clinicians across multiple languages and regions. The scenarios cover a spectrum of acuity: non-acute everyday conversations, high-acuity crises, and complex clinical inquiries.
Related AI News
Enjoyed this? Get more in your inbox.
Weekly AI breakthroughs, tool reviews, and practical guides.