Cloudflare Will Block AI Crawlers From Ad Pages by Default Starting in September
Cloudflare marked the second anniversary of what it calls Content Independence Day on July 1 with its most aggressive move yet in the standoff between publishers and AI companies. Starting September 15, 2026, new customers, new sites from existing customers, and every site on Cloudflare's free tier will default to blocking AI training and agent crawlers on any page that carries advertising, while traditional search crawlers remain allowed. The rules are available to all Cloudflare customers, including those on the free plan.
Three Categories Instead of One Blunt Switch
Cloudflare's original AI bot control, launched last year, offered a simple on-off toggle: block AI crawlers or don't block them. The new system replaces that with three categories based on what a crawler actually does with content once it's collected:
• Search crawlers, which index pages so users can find them, remain allowed by default.
• Agent crawlers, which fetch content on behalf of an AI assistant acting for a user, are blocked by default on ad-supported pages.
• Training crawlers, which collect content to train AI models, are also blocked by default on ad-supported pages.
Cloudflare Product Manager Jin-He Lee and Director of Product Bryan Becker said the goal was to move past a one-size-fits-all block, since website owners want more nuanced options than shutting out all automation at once.
Why Googlebot Gets Caught in the Net
The most consequential part of the update is what happens to crawlers that do more than one job at once. Google, Apple, and Microsoft all currently run bots that mix search indexing with AI training or agent functions in a single crawler. Under Cloudflare's new rules, any mixed-use crawler is evaluated under its most restrictive applicable policy. That means if a site blocks Training crawlers, Googlebot, Applebot, and Bing's bot will all be blocked on ad-monetised pages by default, even though those same bots also handle ordinary search indexing.
Cloudflare has been explicit that this is meant to pressure large search providers into splitting their crawlers apart. The company's own data shows mixed-use crawlers currently account for roughly a third of all crawler activity, and it says it wants that share reduced to zero by mid-2027. Cloudflare has also pointed out that Google's dominant search position gives it access to roughly twice as much web content as other AI companies, since staying discoverable in Google Search effectively requires accepting that the same crawler feeds Google's AI products.
Site owners who don't want Googlebot affected can opt out of the new defaults in their Security settings before September 15, preserving current behaviour for crawlers that handle both training and search.
New Tools: BotBase and Granular Content Use Rules
Alongside the category system, Cloudflare is launching BotBase, a searchable database covering every known bot and AI agent it tracks, giving site owners a single dashboard view of who's crawling their site and how. For Enterprise Bot Management customers, Cloudflare is also introducing content use controls with three tiers: Immediate (no storage or reuse at all), Reference (indexing, excerpts, and links back to the source), and Full (summaries or reproduction of content). These preferences get expressed through an extended version of the Content Signals field in robots.txt. Compliance isn't technically enforced, but Cloudflare says it will track whether Verified Bots follow the declared preferences through BotBase, and bots that ignore them or reproduce content in full risk losing their Verified status entirely.
From Blocking to Getting Paid
Cloudflare is also evolving last year's Pay Per Crawl feature, which let publishers charge AI companies simply for crawling their site, into a new model called Pay Per Use. Under the new approach, publishers are compensated when their content actually appears in an AI-generated answer, not just when a bot happens to fetch the page. Ceramic.ai and You.com are the first announced partners, and Cloudflare separately confirmed integrations with Patreon, which will block training crawlers network-wide, and newsletter platform beehiiv, which is rolling out creator-level AI controls through its own dashboard.
Cloudflare CEO Matthew Prince framed the broader push as a response to the fact that much of the internet's traffic is no longer human, saying the company needs to act faster now that non-human traffic makes up the majority of what moves across the web.
Why This Matters Beyond Publishers
The numbers behind this update explain why Cloudflare is moving now. More than half of AI crawler requests reportedly re-fetch pages that haven't changed, wasting bandwidth for publishers who gain nothing from the repeat visits. Cloudflare also says publishers already block AI crawlers other than Googlebot at close to seven times the rate they block Google's own bot. This gap illustrates just how much leverage Google's search dominance gives it over sites that can't afford to disappear from search results. More than 50 major content licensing agreements between publishers and AI platforms have been signed over the past year, a trend Cloudflare expects its new defaults to accelerate.
What Website Owners Should Do Before September 15
For publishers, bloggers, and businesses running sites on Cloudflare, the practical takeaway is that inaction has a real outcome this time: the defaults are changing regardless of whether a site owner logs in to make any adjustments. A few steps worth taking before the deadline:
• Review current bot management settings and decide whether the new Search/Agent/Training split matches your actual preferences.
• Decide explicitly whether to opt out of the new defaults if you want to preserve current behaviour for mixed-use crawlers like Googlebot.
• Consider whether Pay Per Use or the beehiiv and Patreon integrations are relevant if your content is licensed or monetised through those platforms.
• Enterprise customers should evaluate whether the new Immediate, Reference, or Full content use tiers make sense for different sections of a site.
Conclusion
Cloudflare's latest move pushes the AI-and-publishers fight past simple blocking and into a more structural rework of how crawling, training, and compensation are supposed to work. By tying the new defaults to advertising pages and applying the strictest rule to any crawler that mixes functions, the company has put itself, and by extension the millions of sites it fronts, directly between AI companies and the search giants whose bots have quietly done double duty for years. Whether Google, Apple, and Microsoft actually separate their crawlers in response, or leave Cloudflare's free-tier sites blocking Googlebot by default, will say a lot about how this next round of the AI content fight plays out.


