AI Pulse by Inblix

Anthropic Says Claude Code Auto Mode Catches 89% of Harmful Actions, Beats Humans

TechCrunch AI · Aug 9, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Anthropic Says Claude Code Auto Mode Catches 89% of Harmful Actions, Beats Humans

Starting August 14, Anthropic is flipping the default switch on Claude Code for Pro, Max, and Team accounts. The AI coding assistant will ship in auto mode, meaning it will execute most tasks without stopping to ask for a human sign-off at every step. The only guardrails left in place: actions deemed “irreversible, destructive, or aimed outside your environment” will still trigger a halt.

The company first teased this in March as an experimental feature balancing speed against control. Friday’s announcement makes it the standard experience. The justification is blunt — and backed by a number that’s hard to ignore. In a study with 1,053 paid testers, auto mode caught 89% of harmful actions before they caused damage. When humans had to manually approve each step, that detection rate cratered to 13.6%.

Why the gap? Anthropic points to a very human problem: permission fatigue. The company says users hit “approve” on 97% of prompts in Claude Code, rendering the manual checkpoint almost meaningless. It becomes muscle memory, not a genuine review. Claude Code Head Boris Cherny didn’t mince words about which side he’s on. “The team and I use Auto mode exclusively, and have been for many months,” he posted on X. “I couldn’t imagine going back to permission prompts!”

This is a bet that the machine is a more attentive gatekeeper than the person it’s working for. And it’s paired with new safety scaffolding: prompt injection screening and customizable hard deny rules designed to prevent data exfiltration. The shift raises a question Anthropic seems comfortable answering with data rather than philosophy — if the human in the loop is just rubber-stamping decisions, is the loop actually doing anything? The company’s answer, apparently, is no. So it’s cutting the cord, at least partially, and trusting its model to know when to stop itself.

💡 Key Takeaways

  1. Anthropic's auto mode detected 89% of harmful actions in testing, compared to just 13.6% when humans had to approve each step manually.
  2. The 97% approval rate on manual prompts suggests permission fatigue makes human review a hollow safeguard in practice.
  3. Auto mode will still block actions flagged as irreversible, destructive, or targeting systems outside the user's environment.
  4. New safety features shipping alongside this change include prompt injection screening and customizable hard deny rules.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles