How 262 doctors built HealthBench to test AI's bedside manner
Curated by the Inblix editorial team
Most AI health benchmarks look like medical board exams. Multiple choice. Single turn. A sterile proxy for the messy reality of how people actually use these tools. Anthropic just dropped something different. HealthBench, built with 262 physicians across 60 countries, ditches the Scantron format entirely. Instead, it presents AI models with 5,000 multi-turn, multilingual conversations—simulating everything from a panicked neighbor finding someone unresponsive to a clinician puzzling through a complex case—and then grades the responses using 48,562 custom rubric criteria written by those same physicians.
Each rubric is specific to the conversation. A doctor doesn’t just say “good job” or “needs work.” They outline exactly what an ideal response must include and what it should avoid, weighting each criterion by importance. Think of it as a checklist for competence, tailored to each individual scenario. The grading itself is handled by GPT-4.1, which checks the model’s output against the physician’s rubric. The goal isn’t just to pass a test but to measure whether a response demonstrates the judgment a real doctor would find trustworthy. This moves the evaluation from abstract knowledge recall to applied, context-aware reasoning.
Anthropic deliberately designed the benchmark to be unsaturated—meaning there’s significant room for today’s best models to improve. They didn’t build a test their own models could already ace. That’s the point. A saturated benchmark is a dead end; it tells you how to win a game, not how to get safer or more useful in a clinical setting. By releasing both the benchmark and baseline scores from their own models, Anthropic is setting a public starting line. The pressure is now on the entire field to show measurable progress against criteria that actually reflect physician judgment, not just test-taking ability.
I’m most interested in how this handles adversarial cases. The conversations weren’t just generated synthetically; they included human adversarial testing, where people actively tried to trip the models up. That’s a far cry from clean, curated questions. If a model can handle a user deliberately asking for dangerous advice in a convoluted way and still get a high rubric score, that’s a genuine safety signal. The true test of HealthBench won’t just be whether model scores rise, but whether those rising scores correlate with fewer real-world failures when the stakes are life and death.
💡 Key Takeaways
- HealthBench uses 5,000 multi-turn, multilingual conversations crafted with 262 physicians to evaluate AI, moving far beyond simple medical exam questions.
- Each conversation is graded against a unique, physician-written rubric containing specific, weighted criteria—totaling 48,562 individual checks across the benchmark.
- The benchmark is explicitly designed to be unsaturated, meaning current AI models show substantial room for improvement, incentivizing genuine progress.
- Anthropic has released baseline scores for its own models, establishing a public, transparent foundation for tracking progress in health AI capabilities.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.