AI Pulse by Inblix

AI Detectors Fail Up to 29% of the Time When ChatGPT Mimics Your Style

The Decoder · Jul 19, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AI Detectors Fail Up to 29% of the Time When ChatGPT Mimics Your Style

Here’s a finding that should make every professor and publisher nervous. A new study from Epoch AI shows that while popular AI text detectors catch generic bot-written content almost perfectly, their accuracy craters the moment a language model is asked to imitate a specific person’s writing. We’re talking about a jump from near-zero failure rates to an average of 13 percent of AI-generated texts slipping through completely undetected.

The research put three major detectors—Pangram, GPTZero, and Originality.ai—through their paces against a clean corpus of 495 human-written passages across blogging, fiction, and scientific writing. All human text predates ChatGPT’s November 2022 launch, so there’s zero chance of contamination. For plain vanilla AI prompts, the detectors were surgical: GPTZero and Pangram had zero false positives on human text, and false-negative rates maxed out at just 0.7 percent. But Originality.ai raised eyebrows by incorrectly accusing 19 out of 495 human passages of being machine-made. That’s a false-positive rate of 3.8 percent, and if you’re a student wrongly flagged for cheating, that’s not a rounding error.

The real trouble started when the team fed Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro five samples of a real author’s work and told them to write in that same voice. Suddenly, the detectors’ performance fell off a cliff. Originality.ai missed 18 percent of these style-mimicked texts overall, while Pangram and GPTZero missed 10 and 11 percent respectively. Fiction held up decently, with miss rates between just 1 and 5 percent. But scientific writing is where the whole system breaks down: detectors failed to flag between 24 and 29 percent of style-mimicked academic content. In one extreme case, Pangram let 48 percent of Gemini-generated academic passages walk right through.

Each detector uses fundamentally different methods—Pangram relies on a black-box neural network, GPTZero measures predictability and uniformity in word choice, and Originality.ai hunts for statistical patterns—yet all three share the exact same blind spot. They are highly vulnerable to stylistic imitation, and they all choke hardest on the very genre where detection matters most. An earlier Authors Guild test gave these tools a clean bill of health for not falsely accusing humans. This study completes the picture, and it’s not flattering: just because a detector rarely points the finger at a real person doesn’t mean it’s actually catching the bots.

💡 Key Takeaways

  1. Originality.ai falsely flagged 3.8% of human-written texts as AI-generated, a problem that could undermine trust in automated enforcement before it even scales.
  2. Scientific writing is the worst-case scenario, with detectors missing up to 29% of AI-generated academic text when the model mimics an author's style.
  3. Pangram, GPTZero, and Originality.ai use fundamentally different detection architectures but all exhibit the same vulnerability to style imitation, suggesting this is a systemic weakness rather than a fixable bug.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles