Skip to news

Cloudflare Gives Publishers a Search-and-Training Switch

A new control lets sites stay visible in search while refusing AI training, shifting web governance from robots.txt toward infrastructure.

By THE COLDAI TIMES deskPublished 3 min read517 words

Cloudflare has begun rolling out a new control that lets website owners remain discoverable in search while refusing to let their pages train AI models—a small infrastructure change with potentially large consequences for how the web is financed and indexed. The company announced the setting, called “Disallow AI Training,” on September 15, 2026. [0]

What changed

Cloudflare is replacing its broad “Block AI Bots” switch with separate controls for three kinds of automated traffic: search crawlers, training crawlers and AI agents. The distinction matters because the same crawler has often served multiple purposes. A publisher that blocked a mixed-use crawler could lose search visibility along with AI-training access.

The new setting relies on crawler preferences published through robots.txt, while Cloudflare also offers edge-level controls that can block traffic based on its declared behavior. Cloudflare says Apple, Google and Microsoft have either adopted or committed to honoring the preference, allowing sites to reject training while continuing to appear in conventional search. [0]

The company’s default policy also changed September 15 for new domains and some existing free customers that have not changed their settings: training and agent crawlers are blocked on pages carrying advertising, while search remains allowed. Mixed-purpose crawlers can face the stricter rule when they combine search and training functions. [1]

Why it matters

The move gives publishers a more precise answer to a problem that has become central to the AI economy: how to remain visible without giving away the material that creates traffic, subscriptions or advertising value. Search sends users back to a publisher. Training may create value for a model provider without producing an immediate visit or payment. Separating those uses gives publishers leverage without requiring them to disappear from search entirely.

It also moves the debate from policy statements to network infrastructure. Robots.txt has traditionally been a request based on crawler compliance. Cloudflare’s controls add enforcement at the edge, where traffic can be categorized and blocked before it reaches a publisher’s servers. That could make AI access rules easier to administer, but it also gives a private infrastructure company significant influence over how content moves across the web.

Independent reporting has highlighted the practical difficulty: major crawlers do not always fit neatly into search, training or agent categories, and publishers may struggle to understand which setting affects which service. [1] Cloudflare’s own documentation acknowledges that mixed-purpose bots can be blocked under the most restrictive applicable rule. [2]

What remains uncertain

The system depends on accurate crawler identification and meaningful compliance. A preference cannot solve the problem if an AI company uses undisclosed, mislabeled or rapidly changing crawlers. It is also unclear how much search traffic publishers will retain as answer engines increasingly summarize content without sending users onward.

The bigger test will be economic. If publishers use the controls to deny training access, AI companies may respond with licensing offers, alternative data sources or pressure for broader access. Cloudflare has created a switch; whether that switch becomes a bargaining tool—or merely another setting that sophisticated crawlers route around—will depend on enforcement, transparency and the traffic publishers can still monetize.

Related stories