AI Guardrails Are Backfiring on the Elite Hackers Who Hunt Government Zero-Days
Curated by the Inblix editorial team
The cybersecurity researchers tasked with finding critical zero-day vulnerabilities before criminals do are slamming into AI guardrails designed to stop exactly that. And they’re getting loud about it.
Mark Dowd, a veteran exploit developer who’s spent decades selling unpatched flaws to Western governments, put it bluntly on a recent podcast: “It’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.” He’s not alone. Paolo Stagno, CTO at exploit broker CrowdFense, told TechCrunch that AI companies “essentially treat customers like children who need babysitting” with vetted programs like OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program.
The core frustration isn’t just inconvenience. It’s a fundamental category error. Chris Anley, chief scientist at NCC Group, explained that asking a model to exploit a bug is often the final confirmation that a vulnerability is real and needs patching. A refusal from the AI kills that defensive workflow dead. “‘Fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities,” Anley said. “The two can’t really be unpicked.” It’s a hammer, he argues. You need it to build the house, even though it can also be a weapon.
So what are the pros doing? They’re working around the gatekeepers. Anley’s team often falls back on uncensored open-source models when corporate tools refuse. Stagno avoids cloud-based frontier models entirely for vulnerability discovery, fearing sensitive data could leak or get absorbed into future training. He runs open-source models locally for that. Not everyone is stymied, though. Exploit developer Giuseppe Cali says he doesn’t want AI doing the bug hunting anyway. “I am jealous of my bugs, and I like this game too much to let models play it for me,” he said, using AI only for the grunt work of reverse engineering.
💡 Key Takeaways
- Top exploit developers view AI guardrails as arbitrary corporate overreach, with one calling the vetted access programs patronizing.
- The line between offensive and defensive prompts is often nonexistent, making guardrails that block exploit queries a direct hindrance to patching bugs.
- In response to these restrictions, elite researchers are intentionally routing around cloud AI and using uncensored, locally-run open-source models for sensitive work.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.