OpenAI published MentalHealthBench on September 23, an open benchmark for scoring how AI systems handle mental health conversations, alongside a research paper. The company built it, wrote its scoring criteria with outside experts, and grades every model’s answers using its own system: GPT-5.6 Sol, an automated grader OpenAI trained for this purpose.

That structure matters more than any single result. More than eighty licensed psychologists and psychiatrists took part, drawn from 22 countries. Between them they worked in 19 languages and spanned nearly twenty clinical subspecialties, shaping the benchmark’s scenarios and writing its scoring rubric. But the entity deciding whether a model’s answer satisfies those criteria is GPT-5.6 Sol, a tool OpenAI owns and can update at will. The benchmark is a self-graded exam whose questions came from an expert panel.

The conversations in the benchmark are synthetic, not transcripts of real ChatGPT users. OpenAI built them with privacy-preserving techniques meant to mirror realistic usage, sometimes attaching background details, a recent family loss, for instance, to test whether a model picks up on context instead of answering generically. Scenarios cover adults, teenagers between 13 and 17, caregivers, and clinicians, spread across multiple languages and regions. The benchmark sorts each exchange into one of three severity tiers: ordinary, non-acute conversations at one end; exchanges marked by serious distress that stops short of a true emergency in the middle; and situations demanding immediate real-world intervention at the top.

The grading pipeline carries real rigor even though OpenAI’s own model does the final scoring. Each conversation went through at least three clinical experts, who wrote weighted criteria, from negative ten to positive ten, describing what a good or harmful reply looks like. A criterion only made the final rubric if at least two experts agreed and none contradicted it. Ten behavior categories emerged from that process. They include safety, seeking context, preserving a user’s agency, and giving guidance the user can actually act on. OpenAI describes those categories as improving steadily across newer models, with context-seeking rising in particular, but it did not publish specific scores in a form this benchmark’s charts made readable, so none appear here.

OpenAI ran a separate check on whether the clinical rubric matches what patients themselves value. Forty-four adults across 16 countries, all with prior experience turning to AI when they needed emotional support, rated model responses and wrote their own scoring criteria. That exercise stayed confined to non-acute conversations only. Tone and concrete, actionable steps ranked higher for that group than for the clinical panel, while the clinicians put more weight on gathering context and reading ambiguity. That comparison did not alter how the main benchmark scores anything: the expert rubric stands even where everyday users would grade differently.

OpenAI closes with a quote from Dr. Arthur Evans, the American Psychological Association’s chief executive, who says: “Mental health runs on a continuum from flourishing to everyday stress to acute crisis, and systems engaging people across that range need grounding in clinical science and lived experience.” The company is direct about ChatGPT’s limits too, calling it “not a substitute for therapy or professional care.” Alongside that caveat it lists three safeguards: a Trusted Contact feature, built-in crisis resources, and a separate ChatGPT for Teens experience.

MentalHealthBench is open, so outside labs can run their own models against it and, in principle, build their own graders to check OpenAI’s conclusions. Until someone does, every published result on this benchmark is a vendor grading itself against criteria it also collected: useful groundwork, but not independent verification. Any team citing MentalHealthBench results in a product claim should specify that the scores come from OpenAI’s own grader, not a third party.

Based on OpenAI’s announcement, “Introducing MentalHealthBench,” published September 23, 2026.