GPT-4 Beat Trained Clinicians at Matching Cancer Patients to Trials
Curated by the Inblix editorial team
The quiet bottleneck in clinical trials isn’t a lack of drugs — it’s paperwork. Paradigm, a health tech company focused on trial access, discovered that finding patients eligible for potentially life-saving studies has been a slow, manual grind, with doctors rarely having time to sift through dense medical records. That process creates a biased system where trials fill up with people who simply live near the research site.
Paradigm’s original fix involved building and training specialized healthcare NLP models, one painstaking use case at a time. As one Paradigm executive put it, “You spend a lot of time, and you have to do everything use case by use case.” They assumed a custom medical LLM was the only viable upgrade path. Then they tested GPT-4 against their existing stack, and the results upended their roadmap entirely.
On a blended precision and recall metric, GPT-4 proved at least 10% more accurate than their finely-tuned, domain-specific models. The margin got wider on messy data. “The more complex the information, and the more different places that information sat, the better GPT-4 was,” the team found. In some cases, the general-purpose model even outperformed a team of highly trained human experts, a result Paradigm described as “shocking.”
The downstream effects go beyond just better numbers. Paradigm estimates an eye-popping 90% reduction in the clinician time needed to validate model outputs. A process that once demanded building separate models for every data element can now be handled by GPT-4, with new data extraction tasks going live in days instead of months. The team also pointed to OpenAI’s HIPAA compliance posture as non-negotiable for their provider partnerships. The real test ahead is whether this accuracy translates to their ultimate goal: screening in underserved patients whose medical histories are buried in unstructured notes, and making trial access a function of need, not geography.
💡 Key Takeaways
- GPT-4 outperformed Paradigm's custom-built healthcare ML models by at least 10% on a blended precision/recall metric, and in complex cases beat trained human clinicians.
- Using GPT-4 slashed the need for expert clinician validation time by 90%, freeing up doctors and nurses to focus on actual patient care rather than auditing model results.
- Paradigm can now extract entirely new data elements from medical records in days using GPT-4, a process that previously required months of building and training individual ML models.
- While still being proven, Paradigm believes GPT-4's ability to parse unstructured notes could reduce trial selection bias by accurately screening underserved patients with less standardized medical records.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.