AI Pulse by Inblix

Google's AMIE video AI matched doctors on diagnosis — but only with actors

AI News · Aug 12, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Google's AMIE video AI matched doctors on diagnosis — but only with actors

Google’s AMIE system just cleared a bar that would make any medical resident sweat: in a randomized study with 20 board-certified primary care physicians, the AI matched human doctors on diagnostic accuracy, history-taking, and communication quality during live video consultations. The catch? Every patient was a trained actor reading from a script.

The architecture is the interesting part. AMIE splits video consultations across three separate agents — a talker handling patient dialogue, a planner updating differential diagnoses in the background, and a perception agent watching for non-verbal cues and physical signs. Google’s reasoning: a single model can’t sustain natural conversation speed while simultaneously doing deep clinical reasoning and processing audio-visual input. The talker responds without waiting for every background process to finish, which sidesteps the latency problem that would otherwise destroy patient rapport.

Fifteen actors portrayed conditions across five body systems — cardiopulmonary, abdominal, HEENT, neurological, and musculoskeletal. Evaluators rated AMIE on par with the physician group for management appropriateness and communication quality. The video system actually scored higher than both human doctors and a text-only AMIE baseline on eliciting physical signs and guiding patients through virtual examination maneuvers. Patient actors also preferred video over text chat.

Google is appropriately cautious here. The company explicitly states that studies with real patients and their own health conditions must come before anyone draws conclusions about clinical use. That’s not corporate hedging — it’s the correct posture. OSCE-style evaluations with trained actors are a standard medical education tool, but they measure performance under controlled conditions, not the messy reality of patients who omit key symptoms or describe pain in ways that don’t fit a rubric. Google also built an automated evaluation suite to develop the system before human review, testing perceptual tasks like anatomical laterality and signs of respiratory distress. The automated tests allowed rapid iteration, but Google acknowledges they don’t replicate an end-to-end live video feed. Neither method establishes performance with actual patients, and that gap should shape how health systems think about procurement.

💡 Key Takeaways

  1. AMIE's three-agent architecture separates patient-facing dialogue from slower clinical reasoning, solving the latency problem that would otherwise undermine rapport
  2. Evaluators rated AMIE on par with board-certified physicians on diagnostic accuracy and communication, but every patient was a trained actor following a script
  3. Google explicitly states that studies with real patients must precede any conclusions about clinical deployment, and the automated evaluation methods used for development cannot establish real-world performance

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

← Back to all articles