Cloudflare Adds Separate Controls for Search Crawling and AI Training

Disclosure: HostScore News may include paid PR submissions from third parties. Views expressed belong to the respective companies. Learn more about PR submissions.

Cloudflare has introduced separate controls for search crawling, AI training and user-directed AI agents. The change allows website owners to communicate that their content should not be used for AI training without automatically blocking participating crawlers from traditional search indexing.

The distinction matters as popular AI bots like Googlebot, Bingbot and Applebot perform more than one function. Blocking these mixed-use crawlers entirely can affect both AI access and search visibility.

Cloudflare Separates Three Types of Crawler Activity

Cloudflare now classifies crawler activity under three controls:

  • Search covers crawling used to build a traditional search index.
  • Training covers crawling used to train or fine-tune an AI model.
  • Agent covers an AI agent visiting a page on behalf of a user.

A crawler can fall into more than one category. Googlebot, for example, supports traditional Google Search while Google also operates AI-related services.

Cloudflare introduced a Disallow AI Training setting for this situation. The setting lets an accountable crawler continue accessing the site for search while communicating that the content should not be used for model training.

Apple, Google and Microsoft have provided or committed to mechanisms for respecting separate search and training preferences, according to the Cloudflare announcement (screenshot below).

Full Blocking Can Affect Traditional Search Visibility

Cloudflare’s Block and Block on pages with ads settings now apply to mixed-use crawlers, including Applebot, Bingbot and Googlebot. Selecting either option can affect search crawling as well as AI-related activity.

Website owners who want to remain indexed should not use the full Block setting simply because they object to AI training. Cloudflare recommends using Disallow AI Training when the objective is to preserve search access while opting out of training.

Cloudflare is also replacing the broader “Block AI Bots” label with separate Search, Training and Agent controls. Managed Robots.txt is being replaced by Bot Preference Sync, which communicates the website owner’s selections through supported crawler directives.

Existing Cloudflare Settings Are Being Migrated

Cloudflare says most existing customers do not need to take immediate action because their current preferences will be migrated automatically.

A domain using the previous Block AI setting will generally move to the following configuration:

  • Search remains allowed.
  • AI training is disallowed.
  • Agent access is blocked on pages with advertisements.

Domains that already use Cloudflare’s granular controls will retain the practical effect of their existing selections. Previous Training settings of Block or Block on pages with ads will move to Disallow AI Training.

Website owners can still change these settings after migration. Businesses that rely heavily on organic search should review the resulting configuration instead of assuming that every previous rule transferred as intended.

New Domains Receive Different Recommended Settings

Cloudflare now provides different recommended settings for advertising-supported and non-advertising websites.

Search crawling remains allowed under both configurations. Cloudflare recommends Disallow AI Training and Block on pages with ads for new domains that generate advertising revenue. New domains without advertising are initially offered more permissive Training and Agent settings.

These presets are recommendations rather than permanent restrictions. Customers can change each crawler category during onboarding or through the Cloudflare dashboard.

Crawler Support Still Varies by Operator

Crawler operators use different mechanisms to process publisher preferences. Google supports the Google-Extended robots.txt token for AI training controls. Apple uses Applebot-Extended, while Microsoft currently provides controls such as the NOARCHIVE directive and Bing Webmaster Tools.

Cloudflare says Google’s and Apple’s training opt-outs do not affect traditional search ranking. Microsoft is working toward broader domain-level robots.txt support, which Cloudflare says is targeted for early 2027.

The new Cloudflare controls make crawler management easier, but they do not create a single universal standard. Each operator still controls how it interprets and implements publisher preferences.

HostScore’s Take

Website owners should stop treating all AI crawlers as one category. Search indexing, model training and user-directed agent visits create different benefits and risks. A full crawler block may protect content from one use while removing the site from a valuable discovery channel.

This risk is highest with mixed-use crawlers because one user agent can support several services. For publishers running content websites, you should review the following settings:

  • Cloudflare Search, Training and Agent controls.
  • Existing robots.txt rules.
  • Google-Extended and Applebot-Extended directives.
  • Bing’s available indexing and AI controls.
  • Crawl activity in Cloudflare analytics and server logs.
  • Indexing changes in Google Search Console and Bing Webmaster Tools.

Cloudflare’s change provides more control, but the right configuration depends on the website’s business model. A publisher funded by advertising may restrict AI agents more heavily than a business that wants its documentation cited in AI-generated answers.

The useful approach is to set separate policies for search, training and agent access, then verify the actual crawler behaviour. A dashboard setting alone does not show whether a website remains visible across search engines and AI platforms.

/ Cloudflare Adds Separate Controls for Search Crawling and AI Training

More from HostScore

Submit Your Company News

Looking for publicity opportunities at HostScore.net?

Share your company’s latest achievements, product announcements, and company milestones with our readers. Use this self-service submission form and payment gateway to start instantly.

Submit News (Self-Service)

Explore Our Website

HostScore was established to offer those seeking web hosting solutions the opportunity to learn everything they need to know about hosts – before spending a cent on them