MIT study: AI explanations backfire for doctors but mislead patients the most
Curated by the Inblix editorial team
A new MIT-led study in Nature Medicine exposes a dangerous quirk in medical AI: the very tools designed to make AI trustworthy can make users dangerously overconfident, especially if they don’t have a medical degree. When researchers tested several explainability methods for skin disease diagnosis—including heat maps, similar images, and ChatGPT-style text explanations—the results split sharply along the user’s expertise.
For non-experts, every type of explanation boosted accuracy, but for a troubling reason. These users didn’t become better diagnosticians; they just deferred to the AI. The effect was most pronounced with large language model (LLM) explanations, which often sounded authoritative even when the model was dead wrong. Participants trusted vague, generic text explanations and reported higher confidence in their incorrect answers. “Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output,” said co-author Roxana Daneshjou of Stanford.
Clinicians told a different story. Primary care providers actually performed best when they got just a bare AI prediction with no explanation at all. They didn’t blindly follow incorrect AI reasoning in the same way, using their training to catch bad explanations. But the LLM explanations still delivered the smallest accuracy bump for this group, suggesting that patient-friendly breakdowns just become noise for a trained doctor doing a differential diagnosis.
Lead author Marzyeh Ghassemi framed the core tension as a design problem: “Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error.” The study also found that a fairness-constrained model reduced diagnostic disparities for darker skin tones, but the over-reliance risk remained. As AI diagnosis moves from the clinic to direct-to-consumer search tools, the interface might matter as much as the model itself. A slick text explanation isn’t a safety feature if it just makes errors look more convincing.
💡 Key Takeaways
- Non-experts improved their diagnostic accuracy with AI help not through learning, but by simply deferring to the model—making their performance worse when the AI was wrong.
- LLM-generated text explanations were the most persuasive and misleading format for non-experts, increasing their confidence even in incorrect answers.
- Clinicians performed best with a bare AI prediction and no explanation, showing that the usefulness of an explanation entirely depends on the user's medical expertise.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.