#024
2026-08-05

The Era of Three AI Crawler Classes: Cloudflare Reveals 25% of Bot Traffic Is AI, robots.txt Is No Longer Enough

📰 Want more news?
Browse Full News List

If you think robots.txt still gives you full control over who crawls your site, it's time to rethink that. Cloudflare's latest research data reveals a striking fact: as of May 2026, AI-related crawlers already account for 20.3% of verified bot traffic, and when you add in AI search bots, the figure reaches one-quarter. In other words, one out of every four bot requests is AI-related.

The more critical finding: not all AI crawlers are the same. Blocking them all wholesale could cost you far more than you expect.

Why You Can No Longer Treat AI Crawlers as One Category

In the past, we treated bot traffic as a single entity—either it's a human or it's a bot, block it and move on. But Cloudflare's research points to something far more nuanced: AI crawlers actually fall into three distinctly different classes, each with very different consequences when blocked.

1. Training Crawlers

GPTBot, ClaudeBot, and Google-Extended belong here. They harvest your site's content, which may eventually be used to train large language model weights. Blocking them won't prevent anyone from seeing your pages—it just means your content won't end up in an AI model's "knowledge base." For most businesses, the short-term impact is virtually zero.

2. Retrieval Crawlers

OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot—all sit here. This is where you need to pay attention. These crawlers build the real-time indexes that AI search engines query when generating answers. If they don't scrape your pages, you disappear entirely from cited answers in ChatGPT or Claude. Blocking retrieval crawlers equals voluntarily erasing yourself from AI search results.

3. User-Triggered Crawlers

ChatGPT-User, Claude-User, Perplexity-User. These requests happen when a "real person" explicitly asks an AI to read your specific page—say, a prospective client asking ChatGPT "Look up Lafa System's security solutions for me." If you block these crawlers, you're essentially rejecting customers while they're actively researching you.

The Data: Who Gets Blocked the Most?

Cloudflare's Q1 statistics show that GPTBot is the most-blocked AI crawler via robots.txt, appearing in 5.52% of DISALLOW rules, followed by CCBot at 5.08% and ClaudeBot at 4.88%. But the research also found that roughly 80% of AI crawler traffic comes from training-class bots—meaning most of what gets blocked is actually the category that causes the least direct harm.

The real danger: many security teams configuring WAF or CDN rules don't differentiate among these three classes at all. A simple "block all non-Google search engine crawlers" rule may simultaneously kill your brand's visibility in ChatGPT, Perplexity, and Claude.

CDN-Layer Defense Is Replacing robots.txt

The problem with robots.txt has always existed—it's essentially a "polite request," not an enforcement mechanism. But in the AI era, this weakness becomes far more consequential:

Cloudflare's Next Move: Default Blocking of Training and Agent Crawlers on Sept 15

Cloudflare has announced that starting September 15, 2026, newly registered domains will default to blocking "training-class" and "agent-class" AI crawlers on ad-displaying pages, while search-class crawlers remain allowed. The direction is clear: separate "discoverability in search" from "being used for model training."

For content publishers, this might be good—your articles get cited but not harvested for training. But for B2B service providers, if ChatGPT can't read your solutions page in real time, your lead conversion funnel breaks.

The Invisible Cost: A Real Problem for Small and Medium Businesses

AI crawlers generate large volumes of ineffective traffic, consuming your CDN bandwidth and request quotas. According to Cloudflare's data, if your site gets hit with 100,000 bot requests monthly, roughly 25,000 of them come from AI crawlers. These requests never convert into customers, but they genuinely increase your CDN bills.

The catch: most small and medium businesses lack the resources to distinguish "valuable crawler traffic" from "pure cost-drain traffic." They only see their bill numbers climbing, with no way to filter precisely—so they either block everything (losing search visibility) or allow everything (wasting CDN fees).

The Future: Paid Crawling and Licensing Models

Cloudflare has already launched a "pay-per-crawl" model, letting publishers charge AI companies for access. This means the future won't just be "allow or deny," but moving toward "licensing and pricing." For businesses that need to stay indexed by AI search while controlling costs, this may be the only balanced solution.

💡 LAFA Perspective

With AI crawler traffic making up a quarter of all bot requests, managing them through robots.txt is simply not scalable. Lafa System's AI-driven operations platform identifies all three crawler classes in real time at the CDN layer, intelligently adjusts filtering rules—keeping retrieval crawlers for search visibility while precisely blocking training crawlers that only drain your CDN budget, so every dollar of infrastructure cost actually earns its keep.