You open your hosting dashboard and everything looks fine: solid uptime, fast pages, nothing changed in weeks. But bandwidth creeps up a little every month, PHP workers run busier than before, and a surprise overage line item lands on the next invoice. Your traffic didn't grow — so who's consuming your server? Increasingly, the answer isn't your customers. It's AI crawlers — and they're probably hitting your site right now.
This isn't a hypothetical. Cloudflare's August 2026 data shows fewer than half of all HTML page requests now come from a human; Imperva's annual Bad Bot Report puts automated traffic at 51% of all web activity, with bad bots alone responsible for 37%. The growth rate is accelerating too: TollBit's tracking of publisher networks found the AI-bot-to-human-visitor ratio went from 1:200 at the start of 2025 to 1:31 by Q4 — a 6.5x increase in under a year. Your site is one water drop in this flood.
Regular bots might annoy your cache. AI training crawlers force your server into unpaid labor. Their behavior patterns are what really hurt:
Hosting panels show bandwidth, disk and CPU — but never who is consuming them. A spike on Tuesday afternoon: a successful marketing campaign, or ClaudeBot walking through your entire product catalog? Without bot-level visibility you can't tell — and can't fix it. Many site owners conclude "the site is outgrowing its plan" and upgrade, only to discover they were paying for an army of uninvited crawlers.
The single most important idea in bot defense: stop the bot before it reaches your origin. Once a bot gets to your server, the PHP already ran, the queries already executed, the bandwidth already burned — blocking with a plugin or .htaccess at that point is installing a lock after the door was kicked in. The correct order is identification and blocking at the CDN/WAF edge: the request dies at an edge node, the origin never knows it existed, and it costs you nothing.
What about robots.txt? It's a polite suggestion — well-behaved crawlers like GPTBot and ClaudeBot mostly comply, but impersonators and unscrupulous scrapers ignore it entirely. Real enforcement belongs at the network layer. And "block what, keep what" is a judgment call: pure training crawlers take content and give nothing back — block them. AI search crawlers (OAI-SearchBot, PerplexityBot) can drive referral traffic by citing your pages — worth keeping.
When more than half of your traffic isn't human, an inflating bill is no accident — it's the default. Lafa System's AI ops service detects and blocks invalid bot traffic at the edge, 24/7 — stopping what should be stopped in seconds, keeping what brings real referral value — so your CDN and server spend actually goes to real customers.