The Extended Brief
Have it both ways: stay discoverable in search while disallowing AI training

Brief by The AI News AI newsroom · Sep 15, 2026, 10:12 AM EDT edition
Original reporting by Cloudflare Blog — AI — Bryan Becker · published Sep 15, 2026, 9:00 AM EDT
Site owners can now refuse AI training by major mixed-use crawlers without disappearing from Google, Apple, or Microsoft search results.
Key points
- Cloudflare's Disallow AI Training setting lets sites stay search-indexed while blocking the same crawler from training on their content. source ↗
- Apple, Google, and Microsoft honor the setting or have committed to honor it within a specified time frame. source ↗
- Seventeen percent of Cloudflare sites enable some mechanism to block AI training; fewer than 1% block search bots. source ↗
- By early next year, Cloudflare aims to let sites control how much of their content appears in AI summaries. source ↗
- Cloudflare argues robots.txt alone cannot identify who is crawling, determine why, or stop a crawler that ignores it. source ↗
The data
Fewer than 1% of sites block search bots, while 17% enable some mechanism to block AI training.
Numbers from the original article, machine-verified against its text
Practical applications
- If your site sits behind Cloudflare and depends on search traffic, enable the Disallow AI Training setting instead of blocking mixed-use crawlers outright.
- Audit any robots.txt-only AI opt-outs you rely on, since the file cannot verify crawler identity or intent, and move enforcement to the network layer where possible.
- If you run an ad- or subscription-funded site, revisit your crawler policy now that training refusal no longer costs you search discoverability.
Context
Many large platforms run mixed-use crawlers whose data feeds both search indexing and AI training, so refusing training historically meant losing search visibility. robots.txt is a voluntary text-file convention that cannot verify who is crawling or why, and non-compliant crawlers can ignore it. Cloudflare operates at the network layer, where it says it can identify crawlers and enforce a site's stated preferences.
What to watch
- Whether Apple, Google, and Microsoft demonstrably honor the setting within their committed time frames.
- Cloudflare's promised AI Summaries granularity controls, targeted for early next year, which would extend opt-outs to how much content appears in summaries.
Related briefs
- Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
- Anthropic Makes $13.7 Compute Deal With Trump-Linked Rum Group
- Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
- Inside ‘Project Lily’: The Humans Reading Your ChatGPT Chats
Editorial score 3.8 / 5 · significance 3.5 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Business · Policy & Society
Topics: Copyright · Enterprise AI
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.