Cloudflare Lets Sites Refuse Training Without Leaving Google

Agent Web
Marketing teams rarely opened robots.txt while models trained on their pages. On September 15 Cloudflare gave customers a dashboard setting that refuses that copy and still lets Google, Bing, and Apple index the site.
By Shashi Bellamkonda · September 18, 2026
36.6%
verified crawler traffic that is mixed-use (Cloudflare, 2026)
<1%
Cloudflare sites that block search crawlers
17%
sites that restrict training
2027
Bing robots.txt training preference, targeted
Use Disallow AI Training to keep search and refuse model training. Use Block only if you accept losing Googlebot. Use Allow if you want the training crawlers on the site.

Robots.txt is the text file at the root of a website that tells crawlers which pages they may copy. Content teams and marketing teams spent years ignoring it. Writers produced pages. Search engines were supposed to find them. The file sat in the middle, and almost nobody on the publishing side opened it.

Artificial intelligence training ran through that gap. Models copied articles, product pages, and help text while the people who published the work thought they were only talking to Google. Site owners did not get a notice. Publishers now treat those pages as intellectual property they never licensed, and the controls that decide the next copy are still easy to set wrong.

I have seen the other cost of a wrong setting. A company writes a press release so an announcement can travel, then the page never appears when someone asks an assistant about the news. The writer spent a day on text that did not reach a reader.

On September 15, Cloudflare added a setting that gives customers a way through both problems. Bryan Becker’s engineering post, paired with a San Francisco press note the same day, introduced Disallow AI Training. Search crawlers can keep indexing. Training-only crawlers can be refused. If you want the training bots on the site, you change Training back to Allow. Cloudflare is not inventing robots.txt. Bot Preference Sync writes the dashboard choice into that file so marketing does not have to edit a document it never owned (Cloudflare, 2026).

Search, training, and agents are three different visits

A search crawler copies the page so a results list can send a person back later. That visit can still pay for the site. A training crawler copies the page so a model can learn from it. No person lands, and the writing becomes part of the model. An agent crawler opens the page because a person asked an assistant a question right now, and ads on that load often never meet a human pair of eyes.

Googlebot, Bingbot, and Applebot do search and training under one name. Cloudflare calls those mixed-use crawlers, and they accounted for 36.6 percent of verified crawler traffic on its network, the largest bucket in the September figures. Fewer than 1 percent of Cloudflare sites block search. 17 percent restrict training (Cloudflare, 2026).

The old single switch, Block AI Bots, is being retired. Search, Training, and Agent now sit as separate lines in Security settings. There is still no Disallow option for agents. You allow them, block them on pages that show ads, or block them everywhere.

Disallow keeps Googlebot. Block takes it down with training.

Allow lets crawlers in unless another rule stops them. That is the setting for a documentation set or a how-to library you want inside the next model.

Disallow AI Training publishes a no-training preference in robots.txt. Googlebot, Bingbot, and Applebot can keep indexing for search because Cloudflare lists them as Accountable. Training-only crawlers from Amazon, Anthropic, Meta, and OpenAI are blocked, and Cloudflare says that block does not touch search. Disallow exists only on the Training line. For most publishers this is the match: keep the person who arrives from search, refuse the silent copy into a model.

Block, and Block on pages with ads, now apply to those mixed-use crawlers. Either setting stops Googlebot, Bingbot, and Applebot, so search stops with training. I flagged that risk on July 6, when Cloudflare first named the September 15 default. Disallow is the later switch that refuses training without taking the site out of search.

If you want the announcement to appear when a buyer asks an assistant, keep Search and Agent open on that URL.

Google honors a robots.txt disallow for Google-Extended and says that choice does not change search ranking. Apple honors Applebot-Extended the same way. Cloudflare’s new control writes that kind of preference for you.

Bing still needs a tag on the page until early 2027

Microsoft is in the Accountable group with work still unfinished. Site owners who want Bing to skip training today put a NOARCHIVE robots meta tag on the page, or the matching HTTP header, and can use Block URLs or Content Removal in Bing Webmaster Tools. Microsoft says that tag does not change Bing search ranking. A dedicated no-training line in robots.txt at site or domain level is targeted for early 2027. Until that parser ships, Cloudflare Disallow does not automatically tell Bing to skip training (Cloudflare, 2026).

Krishna Madhavan, who works on Bing’s web data platform, posted the same commitment on LinkedIn the day the Cloudflare blog went up. Search results stay unaffected. The syntax comes later.

Add NOARCHIVE on the pages you do not want in Microsoft training. The Cloudflare switch does not finish that job yet.

New ad sites start with training refused. You can turn it back on.

A domain joined on or after September 15 gets a preset. If the operator says the site earns money from ads, Search is Allow, Training is Disallow AI Training, and Agent is blocked on pages that show ads. Cloudflare’s stated reason is that an ad was placed for a person, training replaces the visit with an answer, and an agent can fetch the page with nobody there to see the unit. Sites that do not run ads start with Allow on all three lines, including training. Either preset can be changed during onboarding or later (Cloudflare, 2026).

A documentation site or a developer blog that wants its how-to pages inside the next model can set Training back to Allow. The tighter default is for ad-supported publishers. Existing zones that already set the granular controls keep the practical effect. Earlier Training choices of Block or Block on pages with ads migrate to Disallow AI Training. Headlines that say millions of sites now block Google overstate what the default does. Open the domain and read the three lines.

Security owns the switch. Marketing feels the missing citation.

In June I said findability often dies in a setting marketing never sees. Security owns Cloudflare. Search performance shows up in a marketing report two months later. I wrote that version on June 11. A quiet press page is the same split. The newsroom posted the URL so a reporter, and later an assistant, could quote the announcement. If Search or Agent is blocked on that page, the model has nothing to fetch when a buyer asks what the company just said.

Markdown for agents, which I covered in August, is the live-answer side of the same distinction. GPTBot copies pages to build a future model. ChatGPT-User arrives because a person asked a question. Anthropic splits those jobs the same way. Keep search and the live fetch if you want the page cited. Refuse the training copy on the pages you treat as inventory.

Cloudflare says summaries in search results are the next control. Each operator still takes a yes or no today. The company wants one screen, by early next year, that sets how much of a page appears in a summary. That work is not part of the September 15 change.

CIO/CTO Viability Question

Sit marketing and security on the same domain. Set Search to Allow and Training to Disallow AI Training unless you want the model to learn those pages. Add NOARCHIVE where Bing traffic matters. Then ask an assistant about your last announcement. If the press page does not appear, the writing did not reach the reader you paid for.

Sources

Cloudflare. “Have it both ways: stay discoverable in search while disallowing AI training.” Cloudflare Blog, 15 Sept. 2026, https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/.

Cloudflare. “Cloudflare Helps End the Search-or-AI-Training Tradeoff.” Cloudflare, 15 Sept. 2026, https://www.cloudflare.com/press/press-releases/2026/cloudflare-helps-end-the-search-or-ai-training-tradeoff/.

Cloudflare. “Your site, your rules: new AI traffic options for all customers.” Cloudflare Blog, 1 July 2026, https://blog.cloudflare.com/content-independence-day-ai-options/.

Bellamkonda, Shashi. “Half Your Website Traffic Is Now Bots. Cloudflare Just Gave You a Deadline to Deal With It.” shashi.co, 6 July 2026, https://www.shashi.co/2026/07/half-your-website-traffic-is-now-bots.html.

Bellamkonda, Shashi. “Stay Findable in 2026: What I Told the Closing Session at Info-Tech LIVE.” shashi.co, 11 June 2026, https://www.shashi.co/2026/06/stay-findable-in-2026-what-i-told.html.

Bellamkonda, Shashi. “Cloudflare Converts Pages to Markdown and Scores Whether Claude and GPT Cite You.” shashi.co, 29 Aug. 2026, https://www.shashi.co/2026/08/cloudflare-adds-markdown-for-agents-and.html.

Disclaimer: This blog reflects my personal views only. Content does not represent the views of my employer, Info-Tech Research Group. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it.