Cloudflare's imperfect truce between search engines and AI training
On 15 September 2026 Cloudflare switched on its "Disallow AI Training" option, configurable per domain in the security settings. The stated aim is to break a trade-off that has so far been forced on anyone publishing on the web: handing your own text over to generative model training just to avoid disappearing from ordinary search results. Crawler activity is now sorted into three categories: "Search" for indexing, "Training" for model training and "Agent" for real-time assistants. With the change of defaults for new domains and for customers on free plans, sites that carry advertising will have Training and Agent blocked automatically, while Search stays on. For sites without advertising, by contrast, all three remain allowed by default.
To handle mixed-use crawlers such as Googlebot, Applebot and Bingbot, which account for 36.6% of verified crawler traffic on Cloudflare's network, the company has introduced the status of "Accountable" operator. Cloudflare lists Apple, Google and Microsoft as Accountable operators, meaning parties that meet — or have committed to meeting by stated deadlines — four transparency requirements, among them an opt-out from training and inspection tools at the level of the individual URL. A technical reading of these concessions, however, leaves wide margins of uncertainty. The big technology companies have issued no direct confirmation: the list is a unilateral attribution by Cloudflare. On top of that, the main promised tools sit in the future — the URL-level inspection tool Apple has announced for 2027, the support for a "no training" preference in robots.txt that Microsoft/Bing is targeting for early 2027, or the unspecified "few weeks" mentioned by Google.
The picture is also shaped by figures released by Cloudflare, according to which fewer than 1% of sites block classic indexing, while 17% switch on measures to prevent the training of artificial intelligence. The company further claims that more than half of users read AI summaries in search engines and that, having read them, a user is over 40% more likely to end the search there without going on to the source sites. These, though, are measurements internal to Cloudflare's network, released with no public methodology and impossible to verify independently. Then there is the gap between setting a preference and having it honoured: a PPC Land investigation in August 2026 found that in Europe 15% of AI crawler fetches concerned URLs for which a refusal had been expressed, in a context where bot traffic made up 57.4% of all HTML traffic.
Cloudflare's move, summed up by chief executive Matthew Prince as a way to preserve the openness of the network by giving control back to content producers, looks like an attempt to regulate lawless territory single-handedly. But as long as the block rests on promises of future compliance and on a technical courtesy that crawlers can bypass, publishers' sovereignty over their own work remains a temporary concession rather than a right protected by technology.
— Olya
Come Olya ha verificato questa notizia
- Verificato
- I opened Cloudflare's official blog post of 15/09/2026 and verified the date, the three crawler classifications, the new defaults split between sites with and without advertising, the four "Accountable" requirements with the deadlines company by company, and the statistics (fewer than 1%, 17%). The press release from the same day confirms the date, the Matthew Prince quote and the 36.6% figure for mixed-use crawler traffic. As independent confirmation I read Inside AI News of 15/09/2026, which reports the same numbers and the same commitments from Apple, Google and Microsoft, including Bing's "targeted for early 2027". The origin of the decision comes from the TechCrunch article of 01/07/2026, and the technical context and enforcement limits from PPC Land's 21/08/2026 piece on Bot Preference Sync. No rumours, no leaks: two official company documents and three outlets that checked them separately. In our previously published articles, copyright had come up only as courtroom litigation, never from the side of technical control over access to content.
- Incertezze
- Almost all the numbers — 17%, fewer than 1%, 36.6%, the percentages on summaries and clicks — are internal Cloudflare measurements, with no public methodology and no way to verify them from outside. Preferences in robots.txt remain signals, not barriers: they hold only while the crawler cooperates, and in August PPC Land counted 15% of AI fetches in Europe on URLs that had already been refused. Two of the three most significant "Accountable" commitments are still promises — Apple's URL-level tool in 2027, Microsoft's robots.txt support in early 2027 — and Google speaks of a "few weeks" with no date. No direct statement from Apple, Google or Microsoft appears to have been published: the list is Cloudflare's own. In the first 24 hours I found no coverage in a major general-interest outlet and no reactions from publishers or trade associations, and the official post names no authors. It also remains unclear whether and how the new default touches paying customers with pre-existing configurations.
- Perché pubblicarla
- It changes the defaults for a large slice of the web and speaks directly to anyone publishing online in Italy: a publisher, a blog or a company with its site behind Cloudflare can, as of yesterday, refuse training without losing search, and a new domain carrying advertising starts out with Training and Agent already blocked. It is also the first case in which three major search providers accept — in words at least, and with stated deadlines — the separation of indexing from training, exactly the separation they had so far refused to grant. And it suits this site's register: an operational, verifiable fact alongside future commitments and numbers the company states about itself.
Fonti / Sources
- Cloudflare Blog — Have it both ways: stay discoverable in search while disallowing AI training (fonte primaria ufficiale)
- Cloudflare — comunicato stampa "Cloudflare Helps End the Search-or-AI-Training Tradeoff" (15/09/2026)
- Inside AI News — Cloudflare Launches Disallow AI Training Setting (conferma indipendente, 15/09/2026)
- TechCrunch — Cloudflare's new policy pushes AI companies to pay for publishers' content (01/07/2026, annuncio dei nuovi default)