Suicide Risk Test Exposes Deadly AI Gap

smartphone displaying a chat interface on chat.openai.com
Photo: Ascannio / Shutterstock

Thirty-five percent of the time, AI chatbots spot that a person is in crisis and then say nothing about where to find real help.

Quick Take

  • A new Scale AI study tested 25 top AI chatbots using 718 realistic crisis conversations written by licensed clinicians.
  • In about 35% of those chats, bots noticed distress but never pointed users to a suicide hotline or other help.
  • Separate research found none of 29 commercial chatbots tested met the bar for an adequate suicide crisis response.
  • OpenAI says it worked with over 170 mental health experts and cut harmful responses by 65 to 80 percent.

What The Scale AI Research Actually Found

Scale AI shared its findings exclusively with TIME on October 9, 2026. Researchers hired 19 licensed clinicians and crisis counselors to write 718 realistic chat scenarios involving mental distress. They then ran those scripts through 25 of the most advanced AI chatbots on the market to see how each one responded.

The results split into two problems. Some bots missed the warning signs entirely. But in roughly 35% of conversations, the chatbot clearly recognized the person was struggling and still failed to mention a crisis hotline or any other real-world resource. Recognizing pain is not the same as acting on it, and that gap is where people fall through.

Why Spotting Trouble Isn’t The Hard Part

Millions of Americans now talk to AI chatbots about anxiety, depression, and worse, often because therapy is expensive or hard to schedule. That makes the referral failure more than a technical glitch. A person reaching out to a chatbot at 2 a.m. may have no other outlet in that moment, and a missed handoff to human help carries real stakes.

A 2026 review in the journal Global Mental Health backs up the pattern. It tested 29 commercial chatbot agents on suicide risk response and found not one met the criteria for an adequate crisis response. The same body of research shows chatbots tend to handle obvious extremes fairly well but stumble badly on the messier, in-between cases that make up most real conversations.

How Bad Does It Get Across The Industry

A separate academic review published in the Journal of the American Medical Informatics Association found that nearly half of the mental health chatbots it examined, 48%, were rated entirely inadequate. Common failures included an inability to give emergency contact information and a lack of understanding of what a user actually meant by their words. That is not a fringe result. It points to an industry-wide design problem, not one bad app.

This fits a broader truth about how these systems get built. Chatbots are trained to keep conversations flowing and users engaged, not necessarily to interrupt that flow with an uncomfortable but necessary referral. Nobody designed them to fail people in crisis. But speed-to-market incentives and engagement-driven design left a safety gap nobody closed in time.

What The Companies Say They’re Doing About It

OpenAI has publicly responded to this exact criticism. The company says it worked with more than 170 mental health experts to help ChatGPT more reliably recognize distress and guide people toward real support, claiming it cut responses that fall short of its own standards by 65 to 80 percent. OpenAI also built its own evaluation tool, MentalHealthBench, with input from more than 80 licensed experts across 22 countries.

Those fixes matter, but they also confirm the problem was real and widespread enough to need a 170-expert overhaul. Parents, lawmakers, and mental health professionals now face a tool millions of people already trust with their darkest moments. Common sense says a product that talks to people about suicide should be required to clear the same safety bar before it ships, not patched after the fact.

Sources:

time.com, ua.news, openai.com, nature.com, cambridge.org, ukr.net