AI Pulse by Inblix

Penda Health’s AI copilot cut clinical errors in a real-world study of 40,000 visits

OpenAI Blog · Jul 13, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Penda Health’s AI copilot cut clinical errors in a real-world study of 40,000 visits

A new study from Penda Health offers a rare look at what happens when an AI copilot gets deployed in the messy reality of frontline primary care—not a lab. Across nearly 40,000 patient visits in 15 clinics around Nairobi, Kenya, clinicians using the organization’s homegrown LLM tool saw a 16% relative reduction in diagnostic errors and a 13% drop in treatment mistakes. The study, conducted in partnership with OpenAI and published with ethical approval from Kenya’s Ministry of Health and AMREF, wasn’t about a model acing a medical exam. It was about whether an unobtrusive piece of software, integrated directly into the electronic health record, could catch human error in real time without annoying the humans it’s supposed to help.

Penda’s copilot, called AI Consult, functions as a background safety net. It doesn’t interrupt a clinician to chat or offer suggestions unprompted. Instead, it sits silently in the workflow, analyzing de-identified documentation as the visit unfolds, and only flags an issue if it spots a potential error—using a simple traffic-light system of green, yellow, or red alerts. The key, according to Penda’s leadership and OpenAI’s analysis, is that the tool was co-developed with the clinicians who would actually use it. This wasn’t a tech demo foisted on a medical staff; the implementation was designed to feel like a second set of eyes, not an extra chore.

This version of AI Consult, which launched in early 2025 and ran on GPT-4o, marks a critical pivot from an earlier, less successful attempt. The first iteration required clinicians to actively request a second opinion from the LLM, a step that broke their concentration and, unsurprisingly, led to low adoption. Dr. Robert Korom, Penda’s Chief Medical Officer, had recognized the potential of LLMs to go beyond rigid, rules-based decision support, but the real victory here is the workflow design. The study’s authors argue this “clinically-aligned implementation,” combined with a serious effort to train staff on the “why” behind the tool, is what closed the gap between a capable model and actual, measurable impact on patient care.

OpenAI is pointing to this research as a template for bridging what it calls the “model-implementation gap”—the stubborn chasm between a model’s theoretical performance and how it’s used in practice. The company noted that its model performance on the HealthBench benchmark doubled from GPT-4o to its newer o3 model, but that raw capability means nothing if doctors won’t use the tool. Primary care, where a single clinician might see infants, pregnant women, and elderly patients with complex chronic diseases all in one shift, is exactly the sort of high-stakes, high-variety environment where errors are both common and preventable. The WHO has said as much. Whether Penda’s results can be replicated in a different health system with a different culture and different IT infrastructure is the next, much harder question.

💡 Key Takeaways

  1. The copilot’s success hinged on passive, background integration into the EHR; the earlier version requiring active clinician requests failed to gain traction.
  2. A 16% relative reduction in diagnostic errors was measured across nearly 40,000 real patient visits, not a simulated benchmark test.
  3. The study was conducted with formal approval from Kenya’s Ministry of Health and an ethics committee, setting a high bar for implementation research.
  4. OpenAI frames this as closing the 'model-implementation gap,' acknowledging that improving benchmark scores like HealthBench is meaningless without real-world clinician adoption.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles