AI Pulse by Inblix

AI reads X-rays with dangerous certainty, new test shows

The Decoder · Jul 19, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AI reads X-rays with dangerous certainty, new test shows

AI models that diagnose from X-rays and MRIs are getting scarily good at being confidently wrong. And that’s a problem no amount of accuracy can fix.

A new benchmark called RadLE 2.0 just put 16 AI models through 200 radiology cases, and the results are a masterclass in misplaced confidence. Human experts scored 988.7 out of 2,000 points. The best AI managed 758. But the headline number hides something more troubling: several models would have ranked higher if they’d simply shut up more often. The scoring system punishes wrong answers delivered with high confidence, while an honest “I don’t know” scores zero but does no damage. This flips the script on how we evaluate medical AI. As one highly cited paper recently argued, benchmarks that only reward accuracy train models to guess. In a clinic, a confident misdiagnosis is a lawsuit and a patient harm event rolled into one.

No single model won across the board. Anthropic’s Claude Fable 5 led on the combined metric for reliability and safety. Google’s Gemini 3 Pro had the best raw accuracy. Meta’s Muse Spark 1.1 was the champ at knowing when to hand a case back to a human, which tracks with Meta’s recent work cutting that model’s hallucination rate nearly in half by having it refuse to answer more often. Then there’s Grok 4.5, which hallucinates significantly more than its predecessor. It knows more, yes, but it’s also far more convinced of its wrong answers. The researchers flagged a clear pattern: open-weight models and medically-specialized models tried to answer almost everything, and they were frequently wrong with full swagger.

This isn’t an academic exercise. Patients are already uploading their MRIs to chatbots and treating the responses as medical advice. A study in npj Digital Medicine showed that popular chatbots routinely give unreliable answers to health questions. Meanwhile, the RadLE 2.0 team directly calls out executives and investors for overhyping what these systems can do. Claims that AI diagnoses better than 99% of doctors are built on anecdotes and simulations, not evidence. As recently as April, a study of 21 supposedly state-of-the-art models found none were ready for unsupervised clinical work. The contrast with two recent studies on autonomous medical AI—MIRA and AMIE, which kept pace with GPs in simulated consultations—is stark. Those papers fueled expectations that AI would soon diagnose independently. RadLE 2.0’s authors push back hard: before a machine makes decisions on its own, it has to know when it shouldn’t.

There’s also the quiet erosion of human skill to worry about. A 2025 Polish observational study found that doctors who regularly use AI during colonoscopies detect significantly fewer precancerous lesions when the tool is turned off. Detection rates fell from 28.4% to 22.4%. The researchers call it the “Google Maps effect”—without the navigation aid, users get lost. Radiology has been here before. In 2016, Geoffrey Hinton famously declared we should stop training radiologists because deep learning was about to replace them. Nearly a decade later, the field is still overburdened, and Hinton walked it back. He’d reduced the profession to pattern matching and missed everything else. The new generation of AI is far more capable, but RadLE 2.0 makes one thing painfully clear: knowing when to speak and when to say “I don’t know” remains a profoundly human skill that these systems haven’t come close to mastering.

💡 Key Takeaways

  1. The RadLE 2.0 benchmark punishes confident wrong answers and rewards saying 'I don't know,' revealing that several AI models would rank higher if they refused to answer more often instead of guessing.
  2. Open-weight and medically-specialized models were the worst offenders, attempting to answer nearly every case and frequently being wrong with high confidence—the exact behavior most dangerous in clinical settings.
  3. A 2025 study found doctors using AI during colonoscopies saw detection rates for precancerous lesions drop from 28.4% to 22.4% when the tool was removed, suggesting a 'Google Maps effect' of skill erosion that compounds the risks of overconfident AI.
  4. Geoffrey Hinton's 2016 prediction that AI would replace radiologists has aged poorly, and the RadLE 2.0 authors argue current claims of superhuman diagnostic ability are built on anecdotes rather than rigorous evidence.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles