Cloudflare gives publishers new ways to block AI traffic

Cloudflare is introducing new controls that let website owners manage AI related traffic by purpose rather than treating every automated visitor the same. Jin-Hee Lee and Bryan Becker announce the changes in an official Cloudflare Blog post, describing separate settings for Search, Agent and Training crawlers.

The distinction matters for publishers and other content businesses. Search crawlers collect pages to answer future queries and may send visitors back to a site. Agents, such as chat based fetch tools or browser automation, visit in real time to complete a task for a user. Training crawlers collect material to train or fine tune AI models.

The new controls are available to all Cloudflare customers, including those on its free tier. They replace the earlier, broader “Block AI Bots” approach with more specific choices. Site owners can therefore allow search indexing while blocking model training, for example.

New defaults for ad-supported pages

Cloudflare says that from September 15, new domains will block Training and Agent traffic by default on pages displaying advertising. Search traffic will remain allowed by default. The company argues that ads indicate a publisher’s aim to attract human visitors, while search can still provide referral traffic.

The policy also affects crawlers with more than one purpose. If a crawler combines Search and Training functions, Cloudflare will apply the stricter applicable setting. This could block bots such as Googlebot, Applebot and BingBot for customers that block Training. Existing customers can opt out before the new defaults take effect.

For Enterprise Bot Management users, Cloudflare is also launching BotBase, a searchable directory of known bots and their classifications. The company plans further controls based on how a bot uses content: immediate interaction without storage, reference use that indexes and links back, or full use that can summarize and reproduce content.

Cloudflare is additionally testing a use preference in robots.txt. The signal does not itself block crawlers, but tells operators how a site owner wants content to be handled. Cloudflare says bots that ignore such preferences risk losing Verified status.

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×