HomeTagsDistressBench

DistressBench

Chatbots Often Spot a Crisis but Don’t Point People to Help, Scale AI’s Own Test Finds

Scale AI built and graded a benchmark of 718 clinician-written crisis conversations. Pooled across 25 models, in 35.3% of conversations where a model fully recognised the crisis, it still failed to route the user towards help. What the company research shows, how it was scored, and its limits.

Top stories

Latest

No posts to display

Latest on TSN

Free TSN tools: crypto calculator, Flux dashboard and more.