51Degrees just published a study of 3 billion real website visits (the year ending 29 May 2026) spanning telecom, media, retail, payments and tech. Three numbers deserve a look from anyone paying a CDN invoice: AI bot traffic is up 62.5% year over year, now represents 13% of all web sessions (up from 8% a year ago), and of the 72 identified AI-related crawlers, AmazonBot and Meta-External-Agent alone account for 63% of that traffic.
For the past year, AI-crawler discourse has been dominated by GPTBot, CCBot and ClaudeBot — the names that made headlines and courtrooms. The 51Degrees data tells a very different story: on a single day, AmazonBot logged 6.43 million sessions (4%+ of all web traffic) and Meta-External-Agent about 4.23 million (3%). Meanwhile ClaudeBot managed roughly 288,000 daily sessions, ranking seventh, and OpenAI Search Bot about 129,000, ninth. The top four crawlers (AmazonBot, Meta, AhrefsBot, BingBot) alone consumed 85% of all crawl sessions.
The load profile matters more than the name. Known Agents describes AmazonBot as making "broad, high-volume sweeps that fetch far more pages per visit than a search crawler" — a fundamentally different shape from a targeted training crawl. It hits your origin's memory, cache hit rate and egress bandwidth directly. Your origin feels the latency; your invoice feels the amount.
The most counterintuitive line in the research: AI crawlers honor robots.txt 95.6% of the time. In other words, there is almost no "disobedient bot" problem — the gap is on the rules side: the average publisher blocks only about 21% of the tracked AI crawlers in its robots.txt. Coverage ranges from 100% down to roughly 2%, and that spread has nothing to do with technical sophistication. It only reflects one question: when did someone last open the file and check it against the current user-agent list?
In one sentence: the bots are following your rules — your rules were written before they existed. A robots.txt from last year, facing the 2026 crawler landscape, is an expired access card.
Each bot session bills against you: origin requests, cache misses, egress bandwidth, database queries — in exchange for value that may be zero. Search crawlers deliver rankings and traffic; that's business. High-volume training sweeps deliver nothing but your bandwidth invoice. So "block or allow" was never a binary: decide bot by bot — allow the ones that drive search referrals, rate-limit or block the high-volume low-value ones, and prioritize the re-fetchers of unchanged pages. And robots.txt is only a polite request; for the rest, rate limiting and WAF rules at the CDN edge are what actually stop the traffic.
Your most expensive bot traffic is usually the one you never knew was there. Lafa System's AI 24/7 log analysis finds your real bandwidth hogs and automatically enforces allow, rate-limit and block policies per bot at the WAF edge — so your CDN bill only pays for the traffic you actually want.