Do AI crawlers respect robots.txt? What each company promises, and what the logs show
OpenAI, Anthropic and Perplexity all say their crawlers obey robots.txt. OpenAI and Perplexity both carve out the fetches a user triggers. Cloudflare's own tests caught one company crawling under a disguised browser identity, and HasData found 234 of 592 sites that ban GPTBot in robots.txt still served it a page. What each company promises, what the evidence shows, whether AI bots are blocked by default, and the layered way to block them.
Short answer
The declared crawlers from OpenAI, Anthropic and Perplexity say they obey robots.txt, but OpenAI and Perplexity exempt fetches a user triggers, and Cloudflare caught undeclared crawling in 2025. Nothing is blocked by default in robots.txt. Use the file for honest bots, edge rules for enforcement, and logins for anything private.
On 8 October 2026 PPC Land published the number every AI-crawler argument will cite this month: of 592 sites whose robots.txt bans GPTBot, 234 still served it a live page.
That figure comes from HasData's AI Crawler Block Index, which PPC Land wrote up and I read on 9 October 2026.
And it is already being quoted as proof that AI crawlers ignore robots.txt.
It proves nothing of the kind. HasData measured whether servers refuse a request carrying the GPTBot name. It did not measure whether GPTBot reads the file and turns around.
Those are two different questions, and most of the confusion about AI crawlers comes from mixing them up.
1. What robots.txt is, according to the people who wrote it
The standard is RFC 9309, which I read on 9 October 2026. Its introduction describes the rules as a request that crawlers are asked to honour, and then says plainly: "These rules are not a form of access authorization."
Google says the same in its introduction to robots.txt, read the same day. The file "is not a mechanism for keeping a web page out of Google", and it is up to each crawler to obey it.
The file is a sign.
A sign works on crawlers that read signs. The rest is about what to put behind it.
2. What OpenAI, Anthropic and Perplexity each promise
I read all three bot pages on 9 October 2026. The pattern repeats: automatic crawlers obey the file, and fetches a person triggers are handled differently.
OpenAI. The bots page lists GPTBot for training and OAI-SearchBot for ChatGPT search, each controlled by its own robots.txt group. OpenAI says a robots.txt change can take about 24 hours to reach its search systems. ChatGPT-User, which visits a page because someone asked about it, is the exception: "robots.txt rules may not apply."
Anthropic. Its support article names ClaudeBot, Claude-User and Claude-SearchBot and says its bots honour robots.txt and support Crawl-delay. Unlike the other two, it treats the user-triggered bot as controllable: disabling Claude-User stops Claude retrieving your page for a user's question. It also warns that blocking its IP addresses instead "may not work correctly or persistently guarantee an opt-out", because the block stops it reading your file.
Perplexity. Its crawler docs say PerplexityBot powers search results and is not used to train models, and that changes can take up to 24 hours to apply. Perplexity-User is the carve-out. Because a user requested the fetch, it "generally ignores robots.txt rules."
So the honest answer to "do AI crawlers respect robots.txt" starts with a split. The background crawlers say yes. Two of the three companies say their user-triggered fetchers may not.
3. What the logs showed when Cloudflare tested it
Promises are documentation. Logs are evidence.
On 4 August 2025 Cloudflare published a report on Perplexity, read on 9 October 2026. It bought fresh domains nobody could have found, disallowed every bot in robots.txt, blocked Perplexity's declared crawlers with firewall rules, then asked Perplexity about the domains.
Perplexity still described the content.
Cloudflare says that once the declared agents were blocked, it saw requests under a generic Chrome-on-Mac user agent from IP addresses outside Perplexity's published ranges. Its own table puts the declared Perplexity-User agent at 20 to 25 million daily requests and the undeclared one at 3 to 6 million, seen across tens of thousands of domains.
Then it ran the same test against ChatGPT. Cloudflare reports that ChatGPT-User fetched robots.txt and stopped when disallowed, and stopped again when shown a block page, with no follow-up from other agents.
Two caveats belong next to that, because this post is meant to be the rigorous one.
First, it is one company's logs and one test, over a year old. Cloudflare itself expected the behaviour to change once the post went live.
Second, Perplexity rejected it. Tech.co, read on 9 October 2026, reports a Perplexity spokesperson calling the report a "publicity stunt" in a statement to The Verge.
We cannot settle that dispute from outside. But the test shows the limit: a crawler willing to change its name walks straight past the file.
4. What the HasData numbers do and do not say
Back to the 234 of 592.
PPC Land is careful about the method. HasData sent one request announcing itself as GPTBot and one as a normal Chrome browser, from the same datacenter address. Its GPTBot requests carried the official name but were not OpenAI's verified crawler, so a server checking real identity would treat them differently.
That makes it a test of servers. A site that bans GPTBot in its file and then serves a page to anything calling itself GPTBot has a sign and no lock. HasData's author put it in four words: "robots.txt is a suggestion".
The same write-up found the reverse too. Some sites say nothing about GPTBot in robots.txt and block it at the firewall anyway, and HasData counted 178 of them in September.
And then the part that should worry anyone who outsourced their file. After Cloudflare's 15 September 2026 changes, HasData's 16 September re-run found 118 sites whose archived robots.txt had carried Cloudflare's managed AI block, and none of them carried it any more. Sites that wrote their own lines kept them.
Your robots.txt is whatever your edge serves today. Check the live URL, not your repository.
5. Are AI crawlers blocked by default?
Not by robots.txt. RFC 9309 says that when no rule matches a URL, the URL is allowed, and that a file with no groups at all applies no rules. A file that never names GPTBot lets GPTBot in.
Defaults come from your infrastructure instead.
Cloudflare's 1 July 2026 post, read on 9 October 2026, announced that from 15 September 2026 new domains would get Training and Agent crawlers blocked by default on pages that display ads, with Search allowed. Its 15 September post changed the detail: a new domain is offered one of two presets, and the ad-supported one publishes a Disallow AI Training line in robots.txt and blocks Agent crawlers only on pages with ads. The preset for sites without ads allows all three.
So "blocked by default" is true only on one CDN, for new domains, mostly on ad-carrying pages. Even there, the training opt-out is a robots.txt line, which is a sign again. We covered what that means for a site with no ads in Is Cloudflare blocking ChatGPT from your website?
6. How to block AI crawlers: three layers
Layer one: robots.txt. It handles every crawler that reads it, which is most of the declared ones. Give each bot its own group and repeat every Disallow inside it, because a named group replaces the wildcard group rather than adding to it. Which bots to allow is a separate decision, covered bot by bot in our AI crawlers and robots.txt guide.
Layer two: the edge. A firewall or CDN rule refuses the request before your server answers. Match the user agent and the source IP together: OpenAI and Perplexity both publish IP lists for each bot, so a request that claims GPTBot from somewhere else can be refused.
Cloudflare's 15 September post makes the case in one line: a robots.txt file cannot identify a crawler "or stop a crawler that ignores it". Mind Anthropic's warning, though. An IP block alone can stop a polite bot reading the file that would have kept it out.
Layer three: what neither layer can do. Neither one stops a crawler that lies about its name and address; Cloudflare caught Perplexity's undeclared traffic by fingerprinting it with machine learning and network signals.
Neither removes content already collected, since Anthropic describes a ClaudeBot block as covering a site's future materials. Neither stops an answer engine building a reply from other sites that quote you, which Cloudflare saw Perplexity do once blocked.
And neither keeps a URL out of Google's index, which needs noindex or a password.
Anything you need private belongs behind a login.
7. What we tell business sites
Most businesses should block nothing.
A clinic, a restaurant or an agency wants to be read by the bots that answer customer questions. NovaTechRay's starting audit is the opposite of blocking: confirm the live robots.txt, the CDN settings and the firewall all let the answering crawlers through.
The VALORAE group's notes on how a new VALORAE page gets found show how much depends on the plumbing nobody looks at.
For a publisher whose content is the product, the three layers are the honest version of blocking. Assume nothing stops every bot.
NovaTechRay will update this post when one of these companies changes its documentation or a new independent test lands.
Per company: OpenAI, Anthropic, Perplexity.
The short version
- Treat robots.txt as a request, because the standard itself says it is not access control.
- Read each company's bot page; user-triggered fetchers are where the promises thin out.
- Check your live robots.txt at its URL, since your CDN may rewrite it.
- Enforce at the edge by matching user agent and published IP ranges together.
- Put anything private behind a login and use noindex to keep pages out of Google.
- Leave the answering crawlers open if you want to be recommended.
This is the work NovaTechRay does at novatechray.com.
Frequently asked questions
Do AI crawlers respect robots.txt?
The declared crawlers mostly say they do. OpenAI's bots page says GPTBot and OAI-SearchBot follow robots.txt but that rules may not apply to ChatGPT-User, which fetches pages a user asks about. Anthropic says all three of its bots honour robots.txt. Perplexity says Perplexity-User generally ignores it. Cloudflare reported on 4 August 2025 that Perplexity also crawled under an undeclared browser identity after being blocked, which Perplexity disputed.
Are AI crawlers like GPTBot blocked by default?
Not by robots.txt. RFC 9309 says a URL is allowed when no rule matches it, so a file that never names GPTBot lets it in. Defaults come from your host or CDN instead. Cloudflare's 15 September 2026 post says new domains are offered presets: an ad-supported site gets a robots.txt Disallow for AI training and Agent crawlers blocked on ad pages, while a site without ads allows all three.
How do I block AI crawlers?
In layers. Add a Disallow group for each crawler's user agent in robots.txt, which the declared bots read. Then add a firewall or CDN rule that matches the user agent together with the IP ranges OpenAI and Perplexity publish, so a request that claims the name but comes from elsewhere is refused. Anything private goes behind a login.
How do I stop AI from crawling my website completely?
You cannot guarantee it on a public site. Google's robots.txt guide says the file cannot enforce crawler behaviour and recommends password protection for anything you need kept from crawlers. Cloudflare's 4 August 2025 report showed a blocked crawler switching to a generic browser identity, which Cloudflare caught with machine learning and network signals, not the file. Authentication is the only complete answer.
Can robots.txt block crawlers?
It can ask them to stay out, and well-behaved crawlers comply. RFC 9309 states the rules are not a form of access authorization, and Google's guide says a disallowed URL can still be indexed if other sites link to it. To keep a page out of Google, use noindex or a password, and remember that a crawler blocked by robots.txt never sees the noindex tag.
Want to know how AI models currently describe your business?
We run a free visibility check across ChatGPT, Perplexity, Claude and Google AI Overviews, then show you exactly which signals are missing.
Book a visibility check