Cloudflare sets deadline to block AI crawlers that bundle search with AI training
NBC News ● Covered by 2 sources
Cloudflare is setting a deadline to block AI bots that quietly mix search crawling with AI training data grabs. It's a jab at crawlers that hide model-training scraping behind a legitimate search-indexing excuse.
Based on reporting by NBC News — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Cloudflare has drawn a line in the sand for AI crawlers that try to have it both ways. The company is moving to block bots that combine search-indexing duties with AI training data collection under one crawler identity, giving site owners no real way to allow one and block the other.
The problem, as Cloudflare frames it, is that some AI companies operate crawlers that fetch pages for both purposes at once. A publisher might be fine with a bot indexing their site for search results but dead against that same traffic being repurposed to train a large language model. When the two functions are bundled into a single crawler, there's no clean opt-out. You either let everything through or block everything, and neither choice matches what most site owners actually want.
Cloudflare sits in front of a huge share of the web, so its crawler policies carry real weight. The company has spent the past year building tools that let publishers control bot access more precisely, including default blocking for unverified AI scrapers and mechanisms for sites to signal exactly what they'll allow. This latest move pushes crawler operators toward transparency: separate your search bot from your training bot, or risk getting shut out entirely once the deadline hits.
For AI companies, this is an operational headache. Splitting crawlers means more infrastructure, more disclosure, and less room to quietly hoover up training data under the cover of search. For publishers, it's a small win in a fight that's been mostly one-sided, where content gets scraped first and permission gets sorted out later, if ever.
My take — AI-written commentary, not fact-checked reporting
Good, honestly overdue. AI companies have been treating 'search crawler' as a convenient disguise for training scrapers, and Cloudflare calling that out forces a separation that should've existed from day one. This is the kind of infrastructure-level pressure that actually changes behavior, way more than another lawsuit or toothless opt-out standard ever will.
Read more about this at: NBC News
Related stories
EU forces Google to share its toys with the other AI and search kids
The Register · 2 months ago ·
30