AI Pulse by Inblix

AI Critics Help Humans Spot 50% More Flaws in Summaries

OpenAI Blog · Jul 18, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AI Critics Help Humans Spot 50% More Flaws in Summaries

Here’s a finding that should matter to anyone worried about AI systems learning to deceive us: when humans are shown AI-generated critiques, they get significantly better at spotting errors they’d otherwise miss. OpenAI’s latest research demonstrates that critique-writing models can boost human flaw detection rates by 50% on model-written summaries, and for deliberately misleading text, the improvement is even starker — humans caught the intended flaw 45% of the time with assistance versus just 27% without it.

The setup was straightforward. Researchers trained language models using supervised learning to write critiques of summaries from short stories and Wikipedia articles. Then they showed labelers eight model-written critiques per summary, with a control group getting no help at all. The summaries came from three sources: AI-written, human-written, and human-written summaries specifically designed to contain subtle, tricky flaws. The assisted group consistently outperformed their unassisted peers.

Scale matters here in a specific way. Larger models proved better at self-critiquing than smaller ones, and they could actually use those critiques to improve their own outputs. Smaller models couldn’t pull off that feedback loop. But there’s an uncomfortable wrinkle: models are better at knowing something is wrong than they are at explaining what’s wrong. That gap between discrimination and articulation didn’t shrink as models got bigger. The researchers flagged this as a priority for alignment work — if a model can detect a flaw but can’t communicate it, we’ve still got a supervision problem.

Let’s not overstate things. Summarization isn’t exactly a Herculean task — humans can evaluate one in about 10 minutes. The real test will be whether this approach holds up on genuinely difficult problems, like auditing a massive codebase or fact-checking a book-length argument. But the direction is promising. The long game, as OpenAI frames it, is building assistants that handle all the cognitive heavy lifting of evaluation so humans can focus on stating what they actually want. Whether we get there without the articulation gap biting us is an open question worth tracking.

💡 Key Takeaways

  1. AI-generated critiques helped human evaluators find 50% more flaws in model-written summaries compared to an unassisted control group.
  2. For deliberately misleading summaries, AI assistance nearly doubled the rate at which humans spotted the intended flaw, from 27% to 45%.
  3. Larger language models are better at critiquing their own outputs and can use those critiques to directly improve results, a capability smaller models lack.
  4. A persistent gap exists between a model's ability to detect a problem and its ability to articulate it clearly — and scaling up models didn't close that gap.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles