Key Takeaways
Mixed-use crawlers now account for more than a third of verified crawler traffic on Cloudflare’s network. These crawlers perform multiple unrelated jobs using a single product token, making it impossible to separate traditional search indexing from AI model training.
The traditional good bot vs bad bot binary framework no longer works because bot identity and purpose are misaligned. A verified bot represents itself truthfully, but verification is a statement about honesty, not business value.
Cloudflare data shows that search engines maintain single-digit crawl-to-refer ratios while training crawlers run into the thousands to one. This disparity means AI models consume heavy resources while rarely sending direct visitors back to your site.
Publishers and site owners must evaluate automated traffic based on cost versus return rather than generalized labels. Read logs by purpose rather than name to determine whether citation value outweighs resource consumption.
A request from Googlebot hits your site. Is it indexing your content for search? Training a model? Fetching an answer for somebody typing into Gemini right now? What if you want to allow indexing but block training?
Right now, you can’t tell what that bot is doing or set blocks based on its intentions. But if it makes you feel better, neither can I, and I oversee the edge network for tens of thousands of sites.
That’s why the good bot/bad bot framework doesn’t actually work: because any request can be good or bad for your business, depending on what you want to allow. That binary thinking oversimplifies a very complex problem.
Why the good bot/bad bot binary began
Even when bot traffic became a hot topic, people were thinking about bots in terms of three buckets, not two.
Bots like Googlebot were “good” because everyone wants to be discoverable on Google. A bot testing credentials on your login page was “bad,” for obvious reasons. Then “gray” bots took up the middle space, taking actions that might be genuinely useful, but could also be genuinely expensive. For example, crawlers that helped provide visibility may also have hammered your origin servers with requests you had to pay for twice: once in usage costs, and again in bounce rates when your site slows down for actual customers.
The good bot/bad bot binary only exists because it was easier for us to wrap our heads around bot traffic with simple terms. People were operating under the assumption that a bot’s identity would tell you its purpose. If you knew the name in the user agent, you’d know what it wants and could sort accordingly. Unfortunately, the reality is much more complex.
Where the binary breaks down
Turns out, that good/bad/gray framework isn’t as useful as we would have liked it to be. That’s because what one person considers “good” might be another person’s “bad.”
Mixed agents use a single product token to perform several unrelated jobs, and you may only want to permit part of their functionality. Googlebot is a well-known case of this pattern. Google’s AI Overviews and AI Mode are both part of Google Search, so Google didn’t give them separate crawler identities. That means I can’t separate the crawl that earns you a traditional search result from the crawl that’s training their models to produce AI Overviews.
That’s the problem with mixed-use crawlers: crawler identity and crawl purpose aren’t consistently aligned across the web. A bot can tell you who it is without clearly telling you why it’s there.
Cloudflare’s data shows mixed-use crawlers now account for more than a third of crawler activity on its network, which is why they’re starting to address the issue. As of September 15, Cloudflare announced it’s replacing the previous “Block AI Bots” capability with more granular controls that allow users to block based on purpose: search, agent, or training. Their goal is to push mixed-use crawler traffic to zero by mid-2027.
Any operators who don’t separate their bots will have them judged by their most aggressive behavior. To be fair to Google, they’ve already earned Cloudflare’s Accountable designation, which they could only have earned by either meeting Cloudflare’s accountability ask or providing a specific timeline for when they intend to meet the new requirements.
Moves like this will start pushing agent operators in the right direction, but it’s clear that we’re still building the language and infrastructure that will allow a website owner to understand what each crawler actually wants.
Until that exists everywhere, you need another way to think about and sort AI traffic that doesn’t rely on every bot announcing its intentions.
Four categories, and “good” isn’t one of them
Here’s how my team looks at traffic sorting outside the good/bad binary.
Malicious. This traffic is so unambiguously hostile that no reasonable site owner would defend it. Credential stuffing, vulnerability probing, the stuff nobody wants on their sites. WP Engine automatically mitigates these at the edge for all of our customers. In Q2 2026, we blocked 34.2 billion malicious requests of this type.
Verified. In these instances, an operator has proven who they are, either through a published IP list with a stable user agent, reverse DNS, or increasingly a cryptographic signature on the request itself. But verification is a statement about honesty, not value. A verified bot represents itself truthfully and doesn’t abuse the access gained by its honesty, but that alone won’t help you decide whether you actually want to give it access.
Likely Automated. It’s not verified, it’s not malicious, and it’s almost definitely not a person. We know it’s a bot, but we don’t know what it wants or who its operator is. This is often the largest bot bucket for most sites.
Likely Human. Human traffic is now the minority, but it’s who you want to prioritize most. After all, they’re the ones who are going to engage with your business.
Verified doesn’t mean valuable
I can tell you a crawler is verified, but I can’t tell you if it’s worth what it costs you in terms of resource usage, visibility, business value, etc. That’s a question about your business, and I don’t run your business.
A publisher whose whole model is ad impressions and a SaaS company that wants to be cited in generated answers could (and likely should) reach opposite conclusions about the exact same training crawler. There is no universal setting, which is precisely why the “good vs bad bot” framework is an oversimplification.
What you can do is make a value judgment. Cloudflare publishes a crawl-to-refer ratio: the number of pages a platform consumes relative to the number of visitors it sends back.
Search engines tend to sit in the single digits, because every crawl is intended to power a direct result to someone’s click. Training and answer crawler ratios, though, can run into the hundreds or thousands to one, because the model answers a user directly and rarely sends a visitor to your actual site.
The ratio alone isn’t necessarily an indictment. Whether you allow bots with high crawl-to-refer ratios comes down to a billing decision.
Sometimes, the bill is worth paying. Being the source an assistant cites has real value even without a click, especially for a human on the other end who is starting their comparison shopping. You want to make their shortlist.
Sometimes the juice isn’t worth the squeeze, but you’re the only one who can make that call.
What to actually do about it
It’s up to you to decide how you handle the bots. Here’s an easy way to think about how you manage automated traffic.
First, stop asking whether a bot is good. Ask what it wants to do, what that will cost you in origin load and bandwidth, and what it provides in return.
Write down what you want before you write a rule. Decide what traffic is worth serving, then configure rules to help you chip away at what’s not.
Read your logs by purpose, not by name. For every crawler consuming a high volume of resources on your site, find out whether the operator separates its agents by purpose. If it does, you have choices to make about what actions to block and what to allow. If it doesn’t, you have exactly one option: Block it or don’t. Your decision comes down to whether blocking access is worth losing whatever it enables, whether that’s a citation, model training, or something else entirely.
Serve all users cached content whenever possible. Serving bots only becomes a problem when they’re requesting lots of deep cuts from your site. Dynamic requests take up more resources and take longer to serve than content that’s cached. On average, dynamic requests on WP Engine’s Advanced Network load in about 325 ms, but when served from Edge Cache, that same content loads in just 11–17 ms.
Writing the most extensive block list possible isn’t going to help you navigate the next few years of uncertainty. What will is a deep examination of your own bot traffic. What’s showing up at your door, what does it want, and is it worthwhile for your business?
Nobody can answer that for you. Not your host, and not the agent operators either. The best we can do on our end is provide you with a little more visibility where it matters most.