Chatbots Often Spot a Crisis but Don’t Point People to Help, Scale AI’s Own Test Finds
Scale AI built and graded a benchmark of 718 clinician-written crisis conversations. Pooled across 25 models, in 35.3% of conversations where a model fully recognised the crisis, it still failed to route the user towards help. What the company research shows, how it was scored, and its limits.
Top stories
Latest
No posts to display
Latest on TSN
Free TSN tools: crypto calculator, Flux dashboard and more.
