How AI’s Confident Lies Endanger Patients

Healthcare professional uses tablet with medical holograms
Photo: metamorworks / Shutterstock

Artificial intelligence can already read an X-ray faster than most doctors, but it still can’t tell when a patient is lying about how much they’ve been drinking.

Quick Take

  • Doctors and researchers say AI tools regularly produce false but convincing medical answers, sometimes called “hallucinations.”
  • Studies show doctors using AI can still be misled by wrong AI suggestions, even when the truth is right in front of them.
  • Major medical groups, including the American College of Physicians and American Medical Association, say AI must stay a helper, not a replacement.
  • AI already helps catch cancer on mammograms and flag dangerous patterns in colonoscopies, showing real value in narrow, defined tasks.

Why Confident-Sounding Answers Aren’t Always Correct Ones

Dr. Ashkan Nasr, a board-certified internal medicine physician, warns that popular AI chatbots “may produce non-factual or fabricated content”. He tells other doctors to check every AI answer against real medical sources before trusting it. That warning matters because these tools write with total confidence, whether they’re right or wrong. A patient can’t tell the difference just by reading the words on the screen.

Dr. Jonathan Chen of Stanford goes further. He says AI “confabulations” have gotten so polished they’re “no longer reliably detectable by prose quality”. In plain terms, the bad answers now sound just as smart as the good ones. Chen also found that in some tests, AI working alone beat doctors who were using AI to help them, suggesting the mix of human and machine doesn’t always add up to better care.

When Doctors Trust the Machine Too Much

Chen’s research points to something called automation bias. That’s the tendency for people to trust a computer’s answer just because it sounds smart and official. He found that 10 to 20 percent of frontier AI model answers could cause harm, often by leaving something out rather than by stating something wrong. A missed warning can hurt a patient just as badly as a false one.

A separate study backs this up in an unsettling way. Researchers found physicians sometimes stuck with a wrong AI suggestion even after they were shown clear evidence proving it wrong. The doctors judged the AI’s advice “unsuitably superficially,” leaning on gut belief instead of digging into the actual data. That’s not a knock on doctors’ intelligence. It’s a warning about how persuasive a slick, fast answer can be, even to trained experts.

Where the Evidence Draws a Hard Line

A large 2025 review looked at how generative AI stacks up against real physicians across many studies. The result: AI performed close to average, non-expert doctors, but it was “significantly inferior to expert physicians,” with a gap of nearly 16 percentage points in accuracy. Translation: AI can pass as a decent generalist. It’s not yet standing shoulder to shoulder with a seasoned specialist.

Emergency rooms make the point even sharper. A study testing AI models against real emergency physicians found the AI “fell short in emulating the complex clinical judgments that physicians make” and concluded the technology is “not ready to replace human expertise in high-stakes settings”. Emergency medicine is exactly where split-second judgment, incomplete information, and gut instinct matter most, and that’s precisely where AI still struggles.

What the Medical Establishment Is Actually Saying

This isn’t just a handful of skeptical doctors talking. The American College of Physicians has stated plainly that AI tools “should enhance human intelligence, not supplant it”. The American Medical Association echoes that stance, insisting AI should work as “an assistive tool,” not “an autonomous decision-maker” in patient care. When two of the largest physician organizations in the country agree on a limit, that’s worth taking seriously.

None of this means AI is useless in medicine. Dr. Leana Wen points out that AI already helps radiologists catch cancer on mammograms and helps flag dangerous colon polyps during colonoscopies. It can also extend basic screening to areas where doctors are scarce. The honest picture isn’t AI versus doctors. It’s AI as a sharp tool in a skilled doctor’s hand, not a replacement for the hand itself.

The Piece Machines Still Can’t Replicate

Dr. Ross Upshur argues that AI can create a false sense of certainty, what he calls “pseudo certainty,” when real medicine often demands sitting with uncertainty and understanding a patient’s “life worlds and values”. That’s the piece no algorithm has cracked. A doctor who notices a patient flinch, who asks the follow-up question a script wouldn’t suggest, who remembers a family history mentioned three visits ago, is doing something no chatbot can fake. Common sense says technology should sharpen that human skill, never replace it.

The lesson for patients and policymakers alike is simple. Push AI into the lab, the scan reader, and the paperwork pile where it saves time and catches patterns humans miss. Keep a licensed, accountable human being making the final call at the bedside. That balance protects patients now, and it keeps the incentive structure honest as these tools keep improving.

Sources:

pmc.ncbi.nlm.nih.gov, pubmed.ncbi.nlm.nih.gov, dl.acm.org, nih.gov, hub.jhu.edu