AI Pulse by Inblix

AI Struggles with Real Knowledge Work, Fails 97% of Tasks

The Decoder · Jun 19, 2026 · 1 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AI Struggles with Real Knowledge Work, Fails 97% of Tasks

A new benchmark, the AA-Briefcase test, evaluates AI models on complex knowledge work projects involving thousands of fragmented source files. Despite top performers like Claude Fable 5, AI models consistently fail to fully solve tasks, with only 3% success rate. This highlights the limitations of current AI technology and underscores the need for more realistic evaluations. Why it matters: This study’s findings underscore the need for more practical and nuanced assessments of AI capabilities, as the current hype surrounding AI may not accurately reflect its true potential.

💡 Key Takeaways

  1. AI models struggle with complex knowledge work projects, failing to fully solve tasks in 97% of cases.
  2. Top performers like Claude Fable 5 achieve high pass rates but still fail to solve 97% of tasks.
  3. The benchmark highlights the limitations of current AI technology and the need for more realistic evaluations.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

← Back to all articles