US Edition
Your source for latest news
TechnologyInternet Infrastructure

Cloudflare begins blocking AI crawlers by default on ad-funded web pages

A new default setting from the internet infrastructure company cuts off bots that mix search indexing with AI training unless site owners opt back in, reshaping how AI firms gather data from millions of websites.

PT
By PressTemps Technology DeskPublished Today, 22:12 ET · 5 min read
Cloudflare begins blocking AI crawlers by default on ad-funded web pages
The entrance area of Cloudflare's San Francisco headquarters. Photo: HaeB / Wikimedia Commons, CC BY-SA 4.0.
What to know
Cloudflare's new default, effective September 15, 2026, blocks AI training and AI-agent crawlers from any ad-monetized page unless the site owner opts in, while continuing to allow conventional search-indexing bots.
The policy applies automatically to all free-tier customers and any new site or domain added to Cloudflare; existing paid customers keep whatever crawler settings they already had configured.
Cloudflare cites its own estimates that it is now roughly 10 times harder to get referral traffic from Google, 750 times harder from OpenAI, and 30,000 times harder from Anthropic than from search engines a decade ago.
A companion "Pay Per Use" program, still in early access with partners including Ceramic.ai and You.com, aims to pay publishers when their content shapes an AI-generated answer rather than merely when it is crawled.

Cloudflare began enforcing a new default policy this week that blocks automated bots collecting data for AI training or AI agents from any web page carrying advertising, unless the site's owner explicitly allows it. The change, which took effect September 15, closes a loophole that let crawlers combining search indexing with AI data collection continue operating freely across much of the web.

The company had announced the policy in July, giving AI companies roughly ten weeks to separate their search-indexing bots from those used to train models or power autonomous agents. Bots that never made that split are now treated as "mixed-use" and blocked by default wherever a page displays ads.

The numbers behind the shift

Cloudflare's argument for the change rests on a widening gap between how much AI systems take from publishers and how much traffic they send back. Chief executive Matthew Prince said it has become roughly ten times harder for a publisher to draw a visit from Google today than a decade ago, roughly 750 times harder from OpenAI's crawlers, and roughly 30,000 times harder from Anthropic's, figures cited in the company's July announcement. The new defaults apply automatically to every new domain onboarding to Cloudflare, every new site created by an existing customer, and all customers on Cloudflare's free tier, according to the company's follow-up post on the September 15 rollout. Paying customers who had already configured their own crawler rules keep those settings; everyone else now defaults to blocking "Training" and "Agent" traffic on monetized pages while continuing to allow "Search" crawlers such as those feeding conventional search results.

How the new rules sort traffic

Cloudflare's documentation splits automated web traffic into three categories: Search, defined as crawling to index content for later retrieval; Agent, meaning a bot acting in real time on behalf of a person; and Training, meaning a crawl used to build or fine-tune a model. A crawler flagged for more than one purpose is now governed by whichever category is most restrictive on a given page, so a bot like Googlebot that Cloudflare classifies as serving both search and AI functions can be blocked on pages where a publisher has restricted Training. Site owners can adjust these settings through a dashboard tool called AI Crawl Control, with no code required.

The categories build on a robots.txt extension Cloudflare introduced last September, the Content Signals Policy, which lets a site declare machine-readable preferences for content it has already been crawled for. In its original announcement of that framework, Cloudflare specified a line such as Content-Signal: search=yes, ai-train=no, covering three separate permissions: whether content may be indexed for search, used as live input to generate an AI answer, or used to train a model outright. At the time, the company said more than 3.8 million domains had already used its managed robots.txt tool to opt out of AI training, well before this week's default change extended similar restrictions automatically.

Who is affected

The practical reach is broad, and much of it falls on publishers with no engineering staff to speak of. Because the block is now a default rather than something a site owner has to switch on, a small blog or local news outlet running on Cloudflare's free plan is automatically covered without ever touching a robots.txt file or a dashboard setting, a deliberate design choice Cloudflare has framed as shifting the burden of asking for access onto crawler operators rather than onto site owners. AI companies, meanwhile, now need explicit permission to keep training on ad-supported pages they could previously reach without restriction. According to TechCrunch's reporting on the rollout, Cloudflare has lined up early partners including the content-licensing platforms Ceramic.ai and You.com to receive payments under a companion program, "Pay Per Use," in which publishers are compensated when their material shapes an AI-generated answer rather than merely when a bot fetches the page.

Large AI operators are not without options. Google and Apple already run separate opt-out crawlers for search versus AI use, which may let them avoid the default block more easily than smaller AI firms still running combined bots. Coverage of the rollout notes that rival infrastructure providers could in theory offer AI companies a way around Cloudflare's restrictions altogether, since the policy only governs traffic passing through Cloudflare's own network.

Reaction

"Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge."

That framing, from Prince, has anchored Cloudflare's public case since July: that automated traffic has overtaken human visitors on much of the web, and that publishers have little practical means of stopping it site by site. Publishers who rely on advertising revenue have broadly welcomed a default that requires opt-in rather than opt-out, since it shifts the technical burden onto crawler operators rather than onto individual site owners who may lack the resources to configure bot rules themselves.

What happens next

The immediate test is how many AI companies choose to separate their crawlers cleanly, negotiate access through Cloudflare's paid programs, or route around the restriction using infrastructure outside Cloudflare's network. Cloudflare has said the Pay Per Use model remains in early access with a limited set of partners, meaning most publishers will not see direct payments soon even as their pages become harder for AI crawlers to reach. Alongside the September 15 changes, Cloudflare also rolled out BotBase, a searchable directory for Enterprise Bot Management customers that catalogues every verified bot and agent Cloudflare tracks, shows 24-hour traffic volumes for each, and lets a security team pull a bot's detection ID to target it directly in firewall rules, giving larger organizations far more granular visibility than the free-tier defaults provide. Whether other infrastructure providers adopt similar defaults, or whether AI companies simply shift more of their data collection to providers with looser rules, will determine how much practical effect the policy has on the economics of AI training data.

More on this story

All Technology