Cloudflare will block AI agents on ad-supported pages by default Sept 15
Curated by the Inblix editorial team
Starting September 15, the open web gets a new gatekeeper. Cloudflare will begin blocking AI agent crawlers by default on any ad-supported page that sits behind its network, a change that lands on every free-tier customer automatically and shifts the ground under the growing fleet of real-time research and shopping bots.
CEO Matthew Prince announced the move on July 1, replacing the company’s single on/off switch for AI bots with a three-category system: Search, Agent, and Training. The real sting is in the defaults. From mid-September, Training and Agent crawlers will be blocked on pages that display ads, while Search crawlers stay allowed. Cloudflare’s logic is refreshingly blunt. An ad is evidence a page was built for a human to land on. A search crawler that sends a reader back is a referral. A bot that scrapes the answer and hands it to someone else in a chat window is not.
This is not a polite robots.txt suggestion. Cloudflare operates at the network level for a huge slice of global traffic, so blocks are real enforcement, not a request. The pages most affected—news, reviews, pricing, product specs—are exactly the pages that enterprise agents depend on. The failure mode for a research agent won’t be a lawsuit. It will be silence, or an answer stitched together from whatever scraps are still reachable. There is a Google-shaped complication too: Googlebot crawls for both search and training as a single entity, so blocking Training takes your search visibility with it. Prince says Cloudflare hopes to “encourage mixed-use crawlers to separate search from agent use and training,” which is a diplomatic way of saying the squeeze is intentional.
Agent builders need to start figuring out which of their systems read as Agent-class—the classification is behavioral, not opt-in. Expect degraded coverage, not clean failure. Publishers have their own homework, including checking whether they’re on a free tier that gets moved automatically, a detail most coverage missed. The announcement also hints at the money side, with players like Ceramic.ai and You.com already paying publishers when content surfaces in AI results. Cloudflare notes that over half of AI crawler traffic is wasted re-fetching unchanged pages, waste both sides should want to price out. The biggest open question is the taxonomy itself. Bots are classified by what the AI companies declare about their own behavior, and a firm that would rather not label its training run as training has an obvious incentive to get creative. For thirty years, web access was free and unlimited. The bill just got itemized.
💡 Key Takeaways
- Cloudflare's September 15 default blocks Agent and Training crawlers on ad-supported pages, hitting free-tier customers automatically and cutting enterprise agents off from the exact content they rely on.
- Blocking Training crawlers also blocks Googlebot entirely because Google doesn't separate its search and training functions, creating a brutal trade-off between content protection and search visibility.
- Cloudflare CEO Matthew Prince explicitly framed the move as pressure to force 'mixed-use crawlers to separate search from agent use and training,' signaling enforcement, not just policy.
- The behavioral classification system relies on AI companies honestly declaring their bots' purpose, with no disclosed mechanism to prevent a training crawler from being labeled as a search bot.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.