Should You Block AI Crawlers With Robots.txt Or At The Server Level?
The short answer is that robots.txt is a polite request and server level blocking is an enforced rule, and that difference is the whole ballgame. If you are trying to decide how to block AI crawlers, the robots.txt vs server level debate really comes down to whether you trust a company like OpenAI or Anthropic to honor a text file you host on your own domain. I have run my own VPS boxes for years and I have seen what happens when a crawler decides the rules do not apply to it. The request gets ignored, the content gets scraped, and you only find out weeks later when someone asks why your product page text shows up in a ChatGPT answer.
This question is surfacing now because site owners are actually looking at their access logs and seeing GPTBot, ClaudeBot, and PerplexityBot hammering their origins. It is no longer a theoretical SEO concern. It is a bandwidth bill and a content theft issue. The source material for this post comes from a Search Engine Journal piece by Helen Pollitt, Head of SEO at Getty Images, who lays out both approaches clearly. I will cover the mechanics of each method, where they fail, and give you a practical audit you can run today to see what is actually hitting your server.
Background: How Robots.txt Blocks Actually Work (And Where They Fail)
A standard GPTBot disallow rule looks like this:
User-agent: GPTBot
Disallow: /
That rule tells OpenAI’s crawler to stay away from every page on your site. If you only want to protect your product pages, you would write:
User-agent: GPTBot
Disallow: /products/
The mechanism is simple and it works exactly the way it is designed to work. The problem is that it only works if the bot is programmed to read the file and obey it. Nothing on your server physically stops GPTBot from fetching your pages. It is a no-trespassing sign in front of an open gate, as Pollitt puts it in the original article. The gate stays open because the server will happily serve content to anyone who asks for it with a valid HTTP request.
The real danger here is not the non-compliant bot. It is the accidental block. If someone on your team makes a mistake and writes:
User-agent: *
Disallow: /
you have just told every search engine, every AI crawler, and every monitoring service to stay away. Google stops indexing you. Bing stops indexing you. Your organic traffic dries up over the course of a few weeks and you have no idea why until someone audits the robots.txt file. I have seen this happen on a client site that was doing 40,000 visits a month. It took them eleven days to notice because nobody was watching the search console alerts.
There is also the maintenance burden. New AI crawlers appear regularly. Every time a new one shows up, someone has to manually add a new disallow rule. That is fine when you have one site, but it becomes a chore when you manage a portfolio of domains. The robots.txt file is not self-updating and it never will be.
What’s Happening Now: The Current AI Crawler Situation
The reputable crawlers that officially support robots.txt today include OpenAI’s GPTBot and OAI-SearchBot, Anthropic’s ClaudeBot, Google’s Google-Extended, and Perplexity’s PerplexityBot. These companies have publicly stated that their crawlers will respect robots.txt rules. That is good news for site owners who want to opt out of training data. You add the disallow rules and you are done.
But here is where I get skeptical. The compliance is voluntary and there is no central body checking whether these companies actually honor the file. Pollitt makes this point in the source article and I agree with her. A company can claim compliance in a blog post and then quietly change their crawler behavior in a software update. You would not know until you check your logs.
Cloudflare and other CDNs now intercept AI bot requests before they hit your origin server. This is a meaningful shift because it saves your server bandwidth. The bot never touches your origin, so your CPU stays idle and your bandwidth bill stays flat. Cloudflare offers preset blocking based on bot type, so you can block all AI training crawlers with a single toggle. That is a huge improvement over editing a text file and hoping for the best.
WAF bot blocking takes this a step further. A Web Application Firewall analyzes request behavior, not just the user agent header. It can catch bots that are spoofing Googlebot or pretending to be a regular browser. This matters because AI crawlers are getting smarter about hiding their identity. A simple header check will miss a bot that claims to be Chrome on Windows. Behavior analysis will catch it because the request pattern does not look like a human browsing your site.
What It Means In Practice: Block AI Crawlers With Robots.txt vs Server Level Rules
So when is robots.txt alone acceptable? If you run a low traffic blog or a personal site and you only care about the reputable crawlers, robots.txt is probably fine. The setup takes five minutes and the maintenance is minimal. You add the four or five disallow rules, you forget about it, and you move on with your life.
When do you need server, CDN, or WAF enforcement? The moment you have content that is actually worth stealing. Product descriptions, proprietary research, detailed how-to guides, pricing pages. If a competitor can scrape your content and use it to train a model that answers questions your customers would otherwise ask you, that is a business problem. Bandwidth costs matter too. I have a client who hosts a large image library and the AI crawlers were pulling tens of gigabytes per week. That is real money on a metered hosting plan.
The practical comparison comes down to three things. Setup effort favors robots.txt because anyone can edit a text file. Maintenance burden favors server level blocking because you configure it once and the CDN or WAF updates its bot database automatically. Blocking strength heavily favors the server stack because it does not depend on the bot’s good faith.
Here is a worked example from my own hosting business. A customer came to me because their site was slow and their bandwidth usage had tripled over a month. I pulled the logs and found that a lesser known AI crawler was hitting their site every few seconds, 24 hours a day. They had no idea who the crawler belonged to and the robots.txt file did not mention it. We added a server level deny rule in Apache using a simple .htaccess entry that matched the user agent string and returned a 403. Traffic stopped immediately and their bandwidth usage dropped back to normal within hours. Robots.txt would not have helped because this particular bot probably did not even read the file.
What To Expect Next: More Bots, More Spoofing, More WAF Dependence
My expectation is that AI crawler traffic will keep climbing for the foreseeable future. Every company that wants to train a model or build a search product needs data, and the easiest data source is the public web. The reputable players will continue to support robots.txt because it keeps regulators off their backs. The smaller players and the outright scrapers will not.
I also expect more bots to ignore or bypass robots.txt entirely. The compliance claims from AI companies will become harder to trust as the financial incentive to scrape grows. You will see more user agent spoofing because it is trivially easy to change a header string. The only reliable defense will be behavior analysis at the WAF level.
This is not a prediction that everything will collapse into a bot war. It is just the direction the traffic data points. Server level and WAF blocking will become the default for any site that treats its content as a business asset. Robots.txt will remain a useful signal for the good actors, but it will stop being a security boundary.
Your Next Step: Audit Your Current AI Bot Blocking Today
Pull your server access logs for the last 30 days and look for every user agent that mentions GPT, Claude, Perplexity, or AI. If you are using Cloudflare, the bot analytics section will show you this without touching the raw logs. On a standard Apache or Nginx setup, a simple grep command against the access log will get you there. Look for repeated hits from the same IP range or the same user agent string.
Add or verify disallow rules for GPTBot, ClaudeBot, Google-Extended, and PerplexityBot in your robots.txt. This takes two minutes and covers the reputable crawlers. If you see repeated hits or any sign of spoofed agents in your logs, create a Cloudflare WAF rule or a server level deny for those bots. The exact rule syntax depends on your stack, but a 403 response is the cleanest way to tell a bot to go away.
Re-check the logs after two weeks. Compare the number of AI crawler requests before and after your changes. If the count dropped to near zero, your blocks are working. If it stayed the same or went up, you have a non-compliant crawler and you need to escalate to a WAF rule that analyzes behavior rather than just the user agent header. That two week check is the only way to know whether you actually solved the problem or just added a sign that nobody reads.