OpenAI launches MentalHealthBench benchmark to evaluate AI on mental health conversations

OpenAI this week released MentalHealthBench benchmark, an open, expert-informed tool designed to evaluate how AI systems respond in realistic mental health conversations. The benchmark, developed with a global cohort of licensed mental health professionals, measures discrete model behaviors such as safety, context-seeking, preservation of user agency, and the provision of actionable guidance where appropriate. OpenAI says the release is intended to let researchers reproduce its methods, run evaluations and extend the work.

MentalHealthBench covers a broad set of conversation types and user personas so evaluations more closely reflect real-world interactions. Scenarios include adults, teens, caregivers and clinicians, and span a spectrum of acuity: everyday non-acute situations, high-acuity conversations that indicate serious concerns, and emergency scenarios that require urgent real-world support. To balance realism with privacy, OpenAI used privacy-preserving techniques to generate synthetic conversations that mirror usage patterns while protecting identifiable information. Some synthetic scenarios include background details—such as recent loss—so evaluators can judge whether a model appropriately uses context to tailor its response.

OpenAI built the benchmark in collaboration with more than 80 licensed psychologists and psychiatrists across 22 countries, covering 19 languages and nearly 20 mental health subspecialties. Each synthetic conversation was reviewed by at least three experts. Those clinicians created detailed rubric criteria for evaluating model replies to the user’s final message; every rubric item targets a single response behavior, for example asking an appropriate follow-up question or offering practical guidance. Criteria carry weights ranging from -10 to +10, with positive weights rewarding beneficial behaviors and negative weights penalizing harmful ones. OpenAI retained rubric items that at least two experts agreed on and that were not contradicted by a third, producing rubrics intended to reflect shared clinical judgment for each scenario.

Responses are scored against these expert-written rubrics using an automated grader, GPT‑5.6 Sol, which OpenAI describes as assessing whether a model demonstrates the ideal behaviors identified by clinicians while avoiding less desirable responses. The benchmark decomposes overall performance into ten interpretable behavioral dimensions for mental health interactions, enabling more granular analysis of strengths and weaknesses across systems. OpenAI reports it evaluated a range of models on the dataset and observed steady improvements among frontier models; it also highlights that models with similar overall scores can differ across the ten dimensions, underscoring areas for targeted improvement.

MentalHealthBench includes conversations involving a teen persona explicitly identified as age 13–17 in system messages. Those teen scenarios were reviewed by clinicians with youth mental health expertise to judge whether responses were appropriate for that age group. OpenAI reiterates that ChatGPT and related tools are not substitutes for therapy or professional care; the benchmark is framed as a measure of progress toward models that can respond with empathy, promote well-being, and guide people toward real-world supports such as localized crisis hotlines or a trusted contact.

Alongside the formal benchmark, OpenAI ran a complementary analysis comparing clinician guidance with what people find helpful in AI support. Forty-four adults who had used AI for mental health or emotional support—representing 16 countries and 14 languages—rated model responses in non-acute conversations and contributed criteria describing helpful support. Participants emphasized practical next steps and tone, while experts placed greater weight on gathering context and careful interpretation of ambiguous situations. OpenAI says this analysis illuminated where user preferences and clinical guidance align and diverge, but it did not change the benchmark’s expert-consensus scoring.

OpenAI released the MentalHealthBench benchmark openly to invite external scrutiny, replication and further research. The organization also cites related efforts such as grants for AI and mental health research, collaborations with external groups including the Partnership on AI, and other complementary evaluation work. The company notes the benchmark accompanies updates to ChatGPT aimed at sensitive conversations, expanded crisis resources, Trusted Contact functionality and a ChatGPT for Teens offering additional protections.

By making MentalHealthBench benchmark public and grounding its design in multidisciplinary clinical expertise, OpenAI aims to provide researchers, developers and clinicians with a shared tool for measuring and improving how AI systems support people’s mental health and safety. The open release is positioned as a starting point for continued evaluation, iteration and collaboration across the AI and mental health communities.

Source: Read the original source

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *