#035
2026-09-11

63% of AI Crawler Traffic Comes From Just Two: AmazonBot & Meta Are Eating Your CDN Bandwidth

📰 Want to read more news?
View full news list

51Degrees just published a study of 3 billion real website visits (the year ending 29 May 2026) spanning telecom, media, retail, payments and tech. Three numbers deserve a look from anyone paying a CDN invoice: AI bot traffic is up 62.5% year over year, now represents 13% of all web sessions (up from 8% a year ago), and of the 72 identified AI-related crawlers, AmazonBot and Meta-External-Agent alone account for 63% of that traffic.

The ones eating your bandwidth aren't the ones in the headlines

For the past year, AI-crawler discourse has been dominated by GPTBot, CCBot and ClaudeBot — the names that made headlines and courtrooms. The 51Degrees data tells a very different story: on a single day, AmazonBot logged 6.43 million sessions (4%+ of all web traffic) and Meta-External-Agent about 4.23 million (3%). Meanwhile ClaudeBot managed roughly 288,000 daily sessions, ranking seventh, and OpenAI Search Bot about 129,000, ninth. The top four crawlers (AmazonBot, Meta, AhrefsBot, BingBot) alone consumed 85% of all crawl sessions.

The load profile matters more than the name. Known Agents describes AmazonBot as making "broad, high-volume sweeps that fetch far more pages per visit than a search crawler" — a fundamentally different shape from a targeted training crawl. It hits your origin's memory, cache hit rate and egress bandwidth directly. Your origin feels the latency; your invoice feels the amount.

The real problem: the bots are well-behaved. The rules just don't exist

The most counterintuitive line in the research: AI crawlers honor robots.txt 95.6% of the time. In other words, there is almost no "disobedient bot" problem — the gap is on the rules side: the average publisher blocks only about 21% of the tracked AI crawlers in its robots.txt. Coverage ranges from 100% down to roughly 2%, and that spread has nothing to do with technical sophistication. It only reflects one question: when did someone last open the file and check it against the current user-agent list?

In one sentence: the bots are following your rules — your rules were written before they existed. A robots.txt from last year, facing the 2026 crawler landscape, is an expired access card.

Every "welcome" has a real price

Each bot session bills against you: origin requests, cache misses, egress bandwidth, database queries — in exchange for value that may be zero. Search crawlers deliver rankings and traffic; that's business. High-volume training sweeps deliver nothing but your bandwidth invoice. So "block or allow" was never a binary: decide bot by bot — allow the ones that drive search referrals, rate-limit or block the high-volume low-value ones, and prioritize the re-fetchers of unchanged pages. And robots.txt is only a polite request; for the rest, rate limiting and WAF rules at the CDN edge are what actually stop the traffic.

Three things you can do this week

💡 LAFA Perspective

Your most expensive bot traffic is usually the one you never knew was there. Lafa System's AI 24/7 log analysis finds your real bandwidth hogs and automatically enforces allow, rate-limit and block policies per bot at the WAF edge — so your CDN bill only pays for the traffic you actually want.