
If you run a blog behind Cloudflare, since September 15, 2026 you can choose between two things that used to be hard to separate: being found in search and still refusing AI training. Cloudflare has reordered its rules for crawlers that do both at once and on September 18 introduced a setting called “Disallow AI Training.” For site owners this sounds like relief. There is a catch, though: choose the wrong option and you may lock Googlebot out.
Key takeaways
- Cloudflare classifies crawlers by purpose: search, training and agent. On September 15, the meaning of the blocking options for crawlers with several purposes was redefined.
- “Block” now also blocks Googlebot, Bingbot and Applebot completely, search included. “Disallow AI Training,” by contrast, prohibits only training and lets search through.
- The old catch-all option “Block AI Bots” is retired, and existing settings are supposed to be migrated automatically.
- For Google, Apple and Microsoft, the training ban rests on their commitment to honor it. For Bing, an automatic training ban is missing until early 2027.
- Site owners should check in their dashboard which setting is actually active rather than trusting the default.
What changed on September 15
Cloudflare set the framework on July 1, 2026, as we reported at the time (in German): crawlers are separated into search, training and agent. “Agent” means real-time requests made on behalf of a person, for example by a chat assistant. Bots with several jobs, so-called mixed-use crawlers such as Googlebot, carry several labels and are treated under all of them. Cloudflare gave providers until September 15 to separate their search and training crawlers.
Since that deadline, according to Cloudflare, “Block” and “Block on pages with ads” also apply to mixed-use crawlers, including Applebot, Bingbot and Googlebot. Anyone who blocks them therefore loses search indexing as well. The catch-all option “Block AI Bots” is being discontinued, and the managed robots.txt is replaced by “Bot Preference Sync.” That mechanism writes the chosen preference into the site’s robots.txt, the text file in which site owners tell crawlers what they want.
New domains get two presets. Without ad revenue, search, training and agent are allowed. With ad revenue, search is allowed, training is prohibited via “Disallow AI Training,” and agents are blocked on pages with ads. For existing domains, Cloudflare carries settings over. A previous “Block AI,” for instance, becomes search allowed, training prohibited and agent blocked on ad pages for domains without granular settings. “In almost all cases, you don’t need to do anything,” Cloudflare writes.
The new option and its price
“Disallow AI Training” is meant to resolve the conflict. The setting prohibits a mixed-use crawler from training but lets it in for search. The precondition is that the operator counts as “Accountable” under Cloudflare’s criteria. It has to meet four requirements or commit to them: an opt-out from AI training via robots.txt or a similar standard, an opt-out for AI summaries, URL-level transparency about which pages are used for training, and an assurance that declining training does not affect classic search results. According to Cloudflare, Apple, Google and Microsoft meet these conditions. Cloudflare blocks pure training crawlers from Amazon, Anthropic, Meta and OpenAI technically, regardless of this.
However, the “Accountable” label is not an external certification but Cloudflare’s own classification, as heise stresses. And the separation itself is not new: Google-Extended, Applebot-Extended and OpenAI’s GPTBot have existed since 2023. What is new is the bundling in the dashboard for the hard case in which a single bot has several jobs.
Where the solution ends
Cloudflare itself names several gaps. Bing will support a “No Training” preference in robots.txt only in early 2027, and until then the setting sends Bing no automatic training ban. Site owners can use the NOARCHIVE meta tag or Bing’s tool for blocking URLs instead. For Apple, a URL-level tool is still missing, and for Google, URL transparency is due “in the coming weeks.” A central control for AI summaries is targeted only for early next year. For agents there is no training-ban equivalent because no established standard exists. A variant only for pages with ads cannot be written into robots.txt.
There is also a fundamental problem, which heise describes this way: robots.txt is not access control but a machine-readable statement of expected behavior. For Googlebot this means the bot fetches the page anyway, and whether the content flows into training depends on Google’s commitment. TechCrunch also points out that Googlebot crawls for AI features such as AI Overviews and AI Mode as well, and that Google points to Google-Extended, which covers only training. Anyone who wants to keep content out of AI answers therefore still has no reliable handle.
Contradictions in the coverage
Anyone who read up beforehand may have run into conflicting statements. An analysis by String Global from August 25 warned that a training block would take Googlebot down with it from September 15, without exception. That matched the state of Cloudflare’s July post. The September 18 post, by contrast, names “Disallow AI Training” as a path that keeps search open. The analysis also reports users who had seen 403 errors for verified Googlebot since July. The sources reviewed contain no confirmation from Cloudflare on this. The scope of the default settings for existing Free accounts is also described differently: TechCrunch named all existing Free customers in July, while Cloudflare’s September 18 post describes a carry-over of existing settings for existing domains.
Conclusion: check the setting, do not guess
The figures Cloudflare cites show how little control has been exercised so far: fewer than one percent of Cloudflare sites block search-engine bots, and 17 percent enable any mechanism against AI training. At the same time, Cloudflare has a commercial interest in this control, for instance in its planned “Pay Per Use” model, meant to have providers pay for content. As Cloudflare describes it, site owners who want to be found but do not want training should pick “Disallow AI Training,” not “Block.” Google Search Console’s crawl statistics remain as a check. The real question, whether providers respect the preferences, cannot be forced by any dashboard setting.
Sources
- Cloudflare Blog: Auffindbar bleiben und gleichzeitig KI-Training untersagen (18.09.2026)
- Cloudflare Blog: Your site, your rules – new AI traffic options for all customers (01.07.2026)
- heise online: Google-Suche ja, Modelltraining nein: Neue Cloudflare-Regeln
- TechCrunch: Cloudflare's new policy pushes AI companies to pay for publishers' content
- String Global: Cloudflare's September 15 Rules Block Googlebot When You Block AI Training

