>
How much is your website spending on CDN every month? If the answer is "a lot," you might be footing the bill for big tech companies — for free.
A recent report from cybersecurity firm DataDome reveals a sobering reality for website operators: Meta's AI bots generated a staggering 9 billion requests in Q2 2026, crawling website content to train their AI models. But these requests delivered zero referral traffic in return — while consuming massive server resources and CDN bandwidth, all paid for by the website owners themselves.
For three decades, an unwritten agreement existed between search engine crawlers and websites: you let me crawl your content, and I send you traffic. Every time Google Bot crawled a page, it earned the chance to appear in search results, bringing real visitors. It was a mutually beneficial ecosystem.
The AI era shattered that rule. AI crawlers scrape content purely for model training — they don't need to — and won't — send visitors back to your site. You pay the CDN bill for them to take your content, and their only return is a larger invoice. As TechRadar's analysis puts it, it's "a unilateral extraction that website owners never agreed to."
Importantly, not all AI companies follow the same approach. The report compares two representative players:
This highlights a critical difference: some AI companies are willing to create value in return, while others simply take without giving back.
In response to this crisis, CDN giant Cloudflare launched a comprehensive AI traffic management upgrade in early July. Following last year's one-click "Block AI Bots" feature, they've introduced the Pay Per Crawl marketplace.
The concept is straightforward: want to crawl website content for AI training? You pay for each request. Cloudflare also introduced a more granular classification system that divides automated traffic into three categories:
Website operators can now set different access permissions for each crawler type — and even charge AI companies for access.
The DataDome report further analyzes how AI crawler traffic impacts websites of different sizes. While large platforms may have negotiating power, small and medium-sized websites and startups bear the heaviest burden:
This problem won't disappear — in fact, it will worsen as more AI companies enter the market. Website operators and service providers should establish three lines of defense now:
First: Identify. Distinguish AI crawler traffic by request characteristics (User-Agent, behavior patterns, IP reputation). Not all automated requests are malicious — Google Bot is still welcome, while training crawlers like Meta's need restrictions.
Second: Classify. Set different rate limits for different crawler types. Legitimate search engines keep their access; model training crawlers should be throttled or blocked.
Third: Block or Charge. Directly block AI crawlers that generate no value, or charge them via mechanisms like Cloudflare's Pay Per Crawl.