Growth & Strategy

Cloudflare sets 15 September deadline to block AI training bots on ad pages

Written by
Full Name
July 16, 2026
From 15 September Cloudflare will block AI training and agent crawlers by default on ad-carrying pages, while leaving search alone — a split that catches mixed-use bots such as Googlebot and hands marketers a decision to make before the deadline.

Cloudflare has given the AI industry a date. From 15 September, the company, which handles traffic for a large share of the web, will block AI training and agent crawlers by default on any page that carries advertising, leaving only search crawlers with automatic access. The controls behind the change went live for every customer, free tier included, on 1 July.

The change matters because the default is flipping from open to closed on the pages that fund most publishing. For B2B marketing teams it lands on the same question that AI Overviews and answer engines already raised: whether their content can still be fetched, summarised and cited when some of the systems doing the summarising are the ones being turned away. It also exposes a bind that has sat under the AI-content debate for two years — that the crawler indexing a page for Google Search is often the same one feeding Google’s models.

What changes on 15 September, and who it affects

Cloudflare will apply the new defaults to every new domain that joins the network, every new site added by an existing customer, and all existing free-tier accounts that have not opted out. Paid customers keep their current settings unless they change them, and any customer can opt out through their security settings at any point before the deadline. On the pages covered, crawlers classified as Training or Agent will be blocked while Search stays allowed.

The mechanism rests on a three-way split the company introduced on 1 July. Cloudflare now sorts AI traffic into Search crawlers, which index a page to answer questions about it later; Agent crawlers, which act in real time on a person’s behalf, such as a chatbot’s fetch bot or a browsing agent; and Training crawlers, which pull content into a model’s weights. An advertisement, in Cloudflare’s reasoning, is a signal that a page was built for a human to land on, so on those pages the training and agent bots are kept out and human attention is treated as the point.

Why Googlebot is the crawler to watch

Cloudflare will judge mixed-use crawlers by all of their behaviours and apply the most restrictive rule that fits. Googlebot, Applebot and BingBot each combine search indexing and AI training in a single crawler, so any site that blocks Training — through the new controls or the legacy “Block AI bots” toggle — will block those bots on its ad pages too. That is the catch for marketers: a site cannot cleanly cut off model training without also risking the search visibility it depends on, unless it opts out and accepts the training that comes with the search crawl.

Cloudflare has aimed the framing squarely at Google, arguing that the search giant has access to roughly twice the information of rival AI companies because staying discoverable has meant accepting AI use as well. By its own figures, mixed-use crawlers that blend search and training now account for 36% of all crawler activity, and publishers already block other AI crawlers at close to seven times the rate they block Googlebot. Google offers a separate opt-out, Google-Extended, that it says lets sites refuse training for products such as Gemini without affecting their place in Search. Chief executive Matthew Prince cast the move as overdue now that “the majority of traffic on the Internet is non-human”.

What marketers should do before the deadline

Marketing teams that run on Cloudflare should start by checking their tier, because free-tier sites are moved to the new defaults automatically and paid sites are not. From there the work is an inventory: which pages carry ads and so fall under the block, and which are the definitive product, pricing and explainer pages an AI system needs to reach to describe a brand accurately. Building that second set as a fetchable, ad-free layer is the practical hedge, since it keeps the material AI answers draw on reachable even where ad-supported editorial is closed off.

The money is the other half of the story. Cloudflare is turning its Pay Per Crawl marketplace into Pay Per Use, which compensates publishers when their content creates value rather than only when a bot fetches it, with Ceramic.ai and You.com as the first partners. The whole taxonomy carries one honest weakness: Search, Agent and Training are categories the AI companies declare about their own bots, and the announcement does not say what stops a firm from labelling a training run as search.

Cloudflare has called the weeks before 15 September a testing window and says the defaults could still change before they take effect. Google, whose Googlebot is the crawler the policy most directly catches, has not said whether it will split the bot in response.

Subscribe to our newsletter

By subscribing you agree to with our Privacy Policy
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Share article

Recommended Reading