Announcing Transluce's Mental Health Evaluation

Transluce | Published: August 31, 2026

Transluce carried out the most expansive independent evaluation to date of how leading AI models respond to users experiencing mental health crises, including suicidal ideation, psychosis, and mania. We outline our findings below, but in summary we found that newer models have significantly improved their responses to users in obvious crisis, while still sometimes struggling with less clear-cut behaviors like creative writing and roleplay about suicide.

Through a collaboration with OpenAI and Anthropic, we received unique, anonymized insights into how real users chat with ChatGPT and Claude for mental health support, which we used to improve the realism of our evaluation. We also received privileged access from OpenAI, Anthropic, and Google DeepMind to study how the APIs that most evaluators test differ from the chatbot apps that most consumers actually use.

We simulated over 50,000 conversations (over 1 million total messages) between users in crisis and 77 model variants released between May 2024 and July 2026 from OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, and Thinking Machines, as well as leading Chinese developers DeepSeek and Moonshot AI.

Studying these sensitive model behaviors in simulation enables us to detect and predict potential risks without exposing real users in crisis. Repeating simulations with the same user profiles also allows us to compare nuanced differences in model behaviors over time, as well as across developers and between model configurations. This type of controlled comparison is impossible in the wild, where no two users have the same conversation twice.

We catalogue our findings in an interactive model behavior report, which shares our measurements as well as the data used to produce them. To enable others to scrutinize and extend our work, we are also releasing:

  1. SimMH-Chat: a dataset of more than 50,000 transcripts and more than 1 million judging results that back our measurements, as well as the synthetic user personas and behavior rubrics we used to generate them, which are also embedded into our interactive technical report.
  2. MHUsage: a privacy-preserving dataset describing how real users chat with ChatGPT and Claude on mental health topics.
  3. Detailed information about the operational conditions of this evaluation and how they may affect the independence, access, and transparency of our results. This includes a template of the legal agreements that we signed with OpenAI, Anthropic, and Google DeepMind as well as a detailed disclosure consistent with the AEF-1 standard from the AI Evaluator Forum.

Below, we briefly summarize key findings across a subset of the models we tested. To view our full results, methodology, examples, and trends over time, see our interactive report.


Structure of our evaluation

To produce our results, we generated multi-turn conversations between AI assistant models and a set of 157 simulated users, many designed to simulate difficult edge cases, and then scored these conversations using automated judges.

Schematic overview of the evaluation: simulated users hold multi-turn conversations with AI assistant models, and automated judges score each conversation for mental-health-related behaviors.
An overview of how we measure model behaviors.

We graded each model’s responses for 14 distinct mental-health-related behaviors, which we defined and validated in partnership with a working group of over 30 clinical experts from a range of organizations including the American Psychological Association, Harvard Medical School, Stanford, and Crisis Text Line. Of these 14 target behaviors, half are clinically validated with a well understood effect on people in crisis, and half cover more exploratory, less clear-cut behaviors that we identified over the course of this evaluation. To validate the reliability of our user simulators and automated judges, we used a range of methods including human labeling, automated diagnostics, and access to anonymized data about production traffic from OpenAI and Anthropic.

“There are many ways an AI system could respond inappropriately, or even harmfully, to someone experiencing suicidal thoughts,” said Dr. Kelly Zuromski, Principal Clinical Research Scientist at Crisis Text Line. “As a suicide researcher, I appreciated the depth and range of model behaviors examined in this work. This kind of comprehensive evaluation is important for understanding where AI may fall short when responding to people in distress.”

What we found

We summarize a few of our key findings below.

1. Testing via API, recent models appear significantly safer than prior-generation models like GPT-4o and Gemini 2.5 which have been associated with tragic cases of suicide and psychosis.1Note that we were not able to test this older generation of models as most consumers would have experienced them, as these models are now only available via developer API and no longer available via their developers’ consumer-facing chatbot apps. Within those apps, these models may have performed differently than in our testing, such as due to differing scaffolding and safeguards. For instance, we found almost no instances in which these newer models endorsed or facilitated suicide, and we found significantly increased rates of helpful behaviors like facilitating connection to human support. The newest models also reinforced delusions and mania much less than earlier versions, in approximately 2-36% of our simulated conversations, depending on the model, versus for example 69%-82% among GPT-4o, Opus 4, and Gemini 2.5.

We also observed that many of the remaining behaviors of concern arise when models engage a user's situation as a practical task to facilitate. For example, recent models remain willing to engage in creative writing even when details suggest it may be about a user’s own suicide, a gray area behavior we identified during our evaluation. The few remaining instances of providing instrumental support for death preparation also tend to take this form (e.g., assistance writing farewell notes to loved ones).

Preview of the interactive report's full table of behavior rates across models.View the data

2. Other than resource banners, models are not generally safer in the consumer-facing chatbot relative to the developer API. We were able to test a series of ChatGPT and Claude models in the browser-based chatbot app (although this did not include older models like GPT-4o or Opus 4 which were no longer available there), as well as a nonpublic Google DeepMind API that is intended to match the Gemini app experience. In the browser, we often triggered popup banners recommending users contact crisis hotlines and similar resources, but due to quirks in how these features are implemented, we could not compare them meaningfully across providers.

A resource banner popup shown in the Claude chatbot app recommending mental health support resources.
An example of a resource banner popup in Claude
A resource banner popup shown in the ChatGPT chatbot app recommending mental health support resources.
An example of a resource banner popup in ChatGPT

Other than these banners, the browser variants we tested largely performed on par with their API counterparts. When they differed, sometimes the consumer-facing app was less safe than the API.

Preview of the interactive report's comparison between API and browser interaction surfaces.View the data

3. In recent models, harmful behaviors tend to be accompanied by helpful behaviors, such as a model reinforcing delusions while simultaneously encouraging users to seek support. Older models more commonly displayed harmful behaviors in the absence of helpful behaviors, rather than both together. However in the most recent generation of models, harmful behaviors are rarer, and when they do occur, they are more likely to occur alongside helpful behavior than on their own.

Preview of the interactive report's analysis of co-occurring helpful and harmful behaviors across models.View our analysis

You can read more details on these findings, including example transcripts and details about our methodology, in our interactive report.

“We are at the very early edge of understanding how these systems will help or hurt human wellbeing. The first step in building a comprehensive response and ensuring AI benefits people is to precisely characterize how these systems behave,” said Anne Maheux, Assistant Professor of Psychology and Neuroscience at UNC Chapel Hill. “A major challenge so far has been knowing what, exactly, to measure. Transluce has developed a detailed framework for characterizing patterns that we expect will matter most, going beyond vague labels like “appropriateness”. The Transluce benchmark offers a foundation for measuring, and eventually understanding and ultimately shaping, how AI affects humankind.”

Validating our simulations against real user behavior

OpenAI and Anthropic granted us unique access to anonymized information about how real users talk about mental-health-relevant topics with ChatGPT and Claude. We did not receive identifying information about users or the contents of their conversations, only data about patterns of user behavior that might impact models’ responses on mental health topics.

We used this data to validate our existing simulators and to create 352 new, production-derived simulated users closer to real usage. For example, the tables below show where we found our original simulators differed most from real user behavior, as well as how we narrowed this gap by producing new simulators derived from this data on production usage.

Bar chart comparing user behavior rates across original simulated users, production-derived simulated users, and real production traffic, for the behaviors most underrepresented in the original simulations.
Largest differences between conversations with real users and with our simulated users, and how creating new simulations based on real conversations helps close the gap.

Validating our original simulators, we find that using these new production-derived simulated users generally does not change how we rank models–e.g., the models with the highest rates of fostering unhealthy user dependency on our original simulators still have the highest rates on the new simulators. This adds an additional source of evidence that our simulations are realistic.

However, we did find that models sometimes exhibit meaningfully different absolute rates of helpful and harmful behaviors when we tested with our production-derived simulators. This adds evidence that better grounding evaluations in real usage patterns can likely help us to move beyond simple benchmarking and to more accurately forecast risks in real-world usage.

Dumbbell chart showing, for several behaviors, the average rate in applicable conversations for original simulated users versus production-derived simulated users, with confidence intervals.
Absolute rates of individual behaviors can shift meaningfully between our original and production-derived simulated users, even when the relative ranking of models is preserved.

For more information on what we learned about users’ mental health usage patterns with ChatGPT and Claude, as well as greater detail on our findings and methods, see our interactive report.

Looking ahead

Rather than trying to set a normative standard for how models should behave in highly nuanced and complex interactions, we instead see this evaluation as part of a descriptive process of uncovering model behavior differences and tracking them across models and over time. It is our hope that this information can feed into a collective, public process to determine how we as a society want models to respond in high-risk scenarios, as part of a broader ecosystem of public oversight, collective sensemaking, and system improvement. To achieve this, it will be essential to improve the realism of behavior evaluations, ground them in the lived experience of actual users, and continue to research the real-world human impacts of the behaviors we measure. We are actively investing in extending this behavior analysis approach to new domains, and see these current results as only a first step.