More people now talk to AI chatbots when they are struggling. A new benchmark from Scale AI, which supplies training and testing data to AI labs, asks whether chatbots respond well when a conversation turns to suicide or self-harm. Its headline finding: models are good at noticing distress, but often stop short of connecting the person with human help.
Spotted via Scale Labs (@ScaleAILabs) on X.
What Scale AI built
Company research. Scale AI designed, ran and graded DistressBench itself. It describes “718 clinician-authored conversations across 23 suicide and self-harm subcategories” [1]. TIME, which first reported the study, says Scale asked “19 licensed clinicians and crisis counselors” to write the chats [3]. “Every conversation is a simulation” [2].
Each conversation comes with a rubric of weighted checks written or revised by clinicians, covering seven areas: recognising the distress, compassion, de-escalation, giving crisis resources, saying the model is not a clinician, avoiding moralising, and avoiding harmful content [2]. A model’s score is the share of that rubric weight its final reply earns. Scale tested 25 frontier models, including models from OpenAI, Anthropic and Google [1][3]. TSN is not listing the rankings.
The 35% figure, exactly
TIME reports that “In about 35% of test conversations in the study, different AI chatbots recognized the user was distressed but didn’t refer them to helpful resources, such as a suicide hotline” [3].
Scale’s paper gives the precise figure. It is neither one model’s result nor an average of models: “pooled across all 25 models, in 35.3% of the conversations where a model satisfied every clinician-written criterion for recognizing the crisis, it failed to route the user toward human help” [2]. “Failed” means the reply missed at least one of the rubric’s crisis-resource checks [2]. That is 12,824 of 36,337 scored conversations, with each model run three times per task [1][2].
The rate varied widely. The strongest model failed to route the user in 15.9% of the conversations it recognised, and the weakest in 69.6% [2]. Longer chats did worse: multi-turn conversations scored 68.5% against 74.2% for single-turn ones [2]. Patrick Oathout, Scale’s Red Team & Safety Lead, told TIME that models “respond in a very just empathetic, kind way as opposed to saying, ‘Okay, it’s time that we get you help'” (reported) [3].
How it was graded
The marking was done by three AI models, not people. Scale picked them by comparing candidates with clinicians’ labels on 75 tasks, and reports agreement of Cohen’s κ = 0.732 [1]. The paper notes that the judges are among the models tested, and that a Google model drafted the first rubrics before clinicians revised them; it reports checks on both [2].
Reported reactions. TIME says an OpenAI spokesperson said the company “did not have sufficient evidence to fully evaluate the Scale AI study”, Google declined to comment and Anthropic did not respond [3].
What this does not prove
- How chatbots behave with real users. These are scripted, clinician-written simulations, scored on the final reply only [2].
- Anything outside English or the US. The benchmark is “English-only, and resource criteria assume U.S. crisis services” [1].
- An independent result. Scale built the test, chose the graders and published the paper. TSN found no sign of peer review.
- Which chatbot is safe. Scale itself says “A high score does not license deployment as a mental-health product” [1].
The Bottom Line
Scale AI’s own benchmark suggests chatbots usually recognise when someone is in crisis but, in 35.3% of those pooled conversations, still fail to point them to human help. It is company research with clear limits, but it measures what refusal tests miss: whether the person was helped.
If you or someone you know needs help
If a life is at risk right now, call your local emergency number (999 in the UK). In the UK you can call Samaritans free, day or night, on 116 123 [4]. In the US, call or text 988 to reach the 988 Suicide & Crisis Lifeline [5]. Elsewhere, contact your local emergency services or a local crisis line.
Related on TSN: Do AI Agents Overstep or Fib? Two New Ways of Checking; AI Still Misreads Pictures, and Reasoning May Hinge on One Word: Two New Studies
Sources
- Scale Labs, “DistressBench” leaderboard and methodology page, and “DistressBench: Evaluating Multi-Dimensional Crisis Support in Large Language Models” summary page, 9 October 2026 (company research). https://labs.scale.com/leaderboard/distressbench ; https://labs.scale.com/papers/distressbench
- Drew Rein, Patrick Oathout, Vishal Kumar and Udari Madhushani Sehwag (Scale AI), “DistressBench: Evaluating Multi-Dimensional Crisis Support in Large Language Models”, paper PDF linked from Scale’s page (company research; not peer reviewed as far as TSN can find). https://static.remotasks.com/uploads/6a8dcf7f9a28fc4d7f8ba553/distressbench-arxiv%20(1).pdf
- Naomi Nix, “AI Chatbots Often Fail To Help Users in Mental-Health Crisis”, TIME, 9 October 2026, 12:00 BST (reported; first report). https://time.com/article/2026/10/09/chatbots-suicide-research/
- Samaritans, “Contact a Samaritan” (checked 9 October 2026). https://www.samaritans.org/how-we-can-help/contact-samaritan/ ; NHS, “Where to get urgent help for mental health” (999 in an emergency). https://www.nhs.uk/nhs-services/mental-health-services/where-to-get-urgent-help-for-mental-health/
- 988 Suicide & Crisis Lifeline (checked 9 October 2026). https://988lifeline.org/
- Scale Labs (@ScaleAILabs), X post announcing DistressBench, 9 October 2026, 15:59 BST (read via the X API). The post itself only announces the benchmark and does not go beyond Scale’s page; a figure circulating from the same thread was not used here. https://x.com/ScaleAILabs/status/2108572962970554460

