ClaudeBot, Claude-User and Claude-SearchBot: what robots.txt does to each of Anthropic's crawlers
Anthropic runs three bots: ClaudeBot for training data, Claude-SearchBot for search indexing and Claude-User for fetches a person asks for. Its help page, read on 9 October 2026, says all three obey robots.txt and Crawl-delay, and it links one IP list for all of them. What a Disallow does to each, the user agent strings, the IP ranges, and why Anthropic warns against blocking by IP.
Short answer
Anthropic runs three bots. ClaudeBot collects possible training data, Claude-SearchBot indexes for search and Claude-User fetches pages a person asks about. Each obeys its own robots.txt line. Anthropic publishes one shared IP list but warns that IP blocks may not hold.
The IP list that Anthropic's crawler help page points to carries a creation date of 7 October 2026. Fetched on 9 October 2026, the file at claude.com/crawling/bots.json held 38 IPv4 prefixes and not one IPv6 range.
That file answers a search people type. Google's autocomplete, read the same day, finishes "claudebot" with "claudebot user agent", "claudebot ip ranges", "claudebot crawler" and "claudebot anthropic".
It also suggests "claudebot vs openclaw" and "clawdbot on mac mini". Neither name appears on Anthropic's page. This post is about the three bots that do.
Here is the complication. I went in expecting to write that Anthropic does not publish its IP ranges. It does.
But it publishes one list for all three bots, and in the same paragraph it tells you not to use it for blocking.
Everything below comes from Anthropic's help article "Does Anthropic crawl data from the web, and how can site owners block the crawler?", fetched 9 October 2026. The old support.anthropic.com address now redirects to support.claude.com, and the page shows a date of 7 April 2026.
Step 1: Three bots, three jobs
Anthropic's page lists exactly three bots.
ClaudeBot collects web content that could contribute to training Anthropic's models.
Claude-SearchBot crawls to improve the quality of Claude's search results.
Claude-User fetches a page when a person asks Claude something that needs it.
Cloudflare's AI Crawl Control bot reference, read 9 October 2026, files them the same way: ClaudeBot as an AI Crawler, Claude-SearchBot as AI Search, Claude-User as an AI Assistant.
So "ClaudeBot" is the name people search, but it is the one of the three with the least to do with whether Claude mentions you.
Step 2: What a Disallow does to each
The page has a column for what happens when you disable each bot. This is the column that matters.
Block ClaudeBot and, per Anthropic, you signal that the site's future material should be excluded from its training datasets. Future, not past.
Block Claude-SearchBot and Anthropic stops indexing your content for search, which the page says may reduce your visibility and accuracy in search results.
Block Claude-User and, in the page's words, it "prevents our system from retrieving your content in response to a user query."
That last one is different from OpenAI's equivalent. OpenAI's bots page, read 9 October 2026, says robots.txt rules may not apply to ChatGPT-User because a user starts the request.
Anthropic makes no such exception. Its page says its bots honour robots.txt, full stop.
So on Claude, the user-initiated line is a real control.
It reaches further than the Claude app, too. Claude Code's public changelog, read 9 October 2026, says version 2.1.83 changed its WebFetch tool to identify as Claude-User. Block that token and a developer asking Claude Code to read your docs page gets refused too.
A correction while we are here. NovaTechRay's AI crawlers and robots.txt guide called ClaudeBot "the general crawler". Anthropic's page describes it as the training collector.
That guide is flagged for a fix, the way the group's correction process says it should be.
Step 3: The robots.txt lines
Anthropic gives two examples. Both use ClaudeBot. Swap the token for either of the other two.
To block one bot from the whole site:
User-agent: ClaudeBot
Disallow: /
To slow it down instead of blocking it, Anthropic supports the non-standard Crawl-delay extension:
User-agent: ClaudeBot
Crawl-delay: 1
Two details people miss.
First, the file goes in your top-level directory, and Anthropic says to repeat it for every subdomain you want to opt out. A rule on www does nothing for shop or blog subdomains.
Second, each token is separate. A ClaudeBot group does not cover Claude-SearchBot. If you want all three treated the same, write three groups.
For a business that wants to be cited, the sensible split is the one in our Cloudflare AI crawler decision tree: allow Claude-SearchBot and Claude-User, and decide ClaudeBot as a separate, lower-stakes training question.
Step 4: The user agent string nobody agrees on
This is where the "claudebot user agent" search goes wrong.
Anthropic's page names the tokens. It does not print the full header string any of them sends.
So people copy strings from directories, and the directories differ. PPC Land's article on AI crawler strings, published 28 June 2026 and read 9 October 2026, gives ClaudeBot as a short string with no AppleWebKit part. Known Agents, the directory at darkvisitors.com, which shows its insights as last updated 8 October 2026, gives a longer string that includes AppleWebKit.
They disagree on Claude-SearchBot too. PPC Land's version ends in a URL, anthropic.com/claudesearchbot.
When I requested that URL on 9 October 2026 it returned a 404. Known Agents' version ends in an email address instead.
I cannot tell you which one is current. Neither can a directory, unless it is reading its own logs.
So do not match the full string. Match the token. "ClaudeBot/", "Claude-SearchBot/" and "Claude-User/" appear in every version I found, and they are what robots.txt keys on anyway.
Step 5: The IP ranges, and the warning beside them
Anthropic's page links its list in one sentence: a crawler whose source IP is on the list is coming from Anthropic.
The file, fetched 9 October 2026, carries a creationTime of 7 October 2026. It holds 38 IPv4 prefixes and no IPv6 prefixes. And it does not say which bot uses which range.
Compare OpenAI. Its bots page, read the same day, links four separate files: one each for GPTBot, OAI-SearchBot, ChatGPT-User and its ad bot. With Anthropic you get one file for everything.
That changes what the list is good for. It tells you a request came from Anthropic, but not which of the three bots sent it. The user agent token does that.
So verification is a pair: the token says which bot it claims to be, and the IP says whether the claim came from Anthropic. Anything with "ClaudeBot/" in the header and an IP outside the list is somebody else using the name.
Then the warning. Anthropic says blocking its IP addresses may not work correctly or persistently guarantee an opt-out, because it stops the bots reading your robots.txt.
Do not block by IP.
It sounds backwards, so here is the reasoning. If the bot cannot fetch robots.txt, it never sees your Disallow.
And a list dated 7 October can change again, so a firewall rule copied today drifts out of date on its own. robots.txt is the opt-out Anthropic says it honours. The IP list is for checking logs.
Anthropic also says its bots will not try to bypass CAPTCHAs. If a bot-challenge service sits in front of your site, a challenged Claude-SearchBot request is a page Claude did not read.
Step 6: What our own robots.txt says
Our robots.txt, read 9 October 2026, allows all three tokens and cites Anthropic's help article as the source for the spellings. The header says it was reviewed on 10 September.
It also had a mistake. The comment above the user-initiated block said those lines were a statement of intent, not a control, because ChatGPT-User and Perplexity-User document that they ignore robots.txt. Claude-User sits in that same block.
For Claude-User, that comment was wrong. Anthropic's page says its bots respect robots.txt, so that line does something. NovaTechRay keeps it set to Allow, and we rewrote the comment on 9 October 2026 so nobody reads it as decoration.
We did not catch it until we read the page line by line for this post. That is the case for rereading operator docs on a date, every time, and writing the date down.
The same breakdown for OpenAI and Perplexity is in GPTBot, OAI-SearchBot and ChatGPT-User and PerplexityBot and Perplexity-User.
The short version
- Write three robots.txt groups for Anthropic: ClaudeBot, Claude-SearchBot and Claude-User.
- Treat ClaudeBot as the training decision and the other two as the visibility decision.
- Remember that a Claude-User block also refuses Claude Code's WebFetch.
- Repeat the rules on every subdomain you want covered.
- Match the token in your logs, never the full user agent string.
- Use claude.com/crawling/bots.json to verify requests, and keep blocking in robots.txt.
NovaTechRay works on getting businesses fetched, read and cited by AI assistants, at novatechray.com.
Frequently asked questions
What is ClaudeBot?
ClaudeBot is one of three bots Anthropic lists on its help page about web crawling, read on 9 October 2026. Anthropic says it collects web content that could contribute to training its models. The other two are Claude-SearchBot, which indexes pages for search, and Claude-User, which fetches pages when a person asks Claude a question.
What is the ClaudeBot user agent?
Anthropic's help page names the robots.txt token, ClaudeBot, but does not print the full header. Two directories read on 9 October 2026 disagree on it: PPC Land lists a short string with no AppleWebKit part, and Known Agents (darkvisitors.com) lists one with it. Both contain ClaudeBot/1.0, so match on that token rather than the whole string.
Does Anthropic publish ClaudeBot IP ranges?
Yes. Anthropic's help page links a list at claude.com/crawling/bots.json. Fetched on 9 October 2026 it was dated 7 October 2026 and held 38 IPv4 prefixes, no IPv6, and no label saying which bot uses which range. Anthropic says the list is for confirming a crawler is theirs, and warns that blocking by IP may not work or last.
How do I block ClaudeBot with robots.txt?
Add a group with User-agent: ClaudeBot and Disallow: / to the robots.txt in your top-level directory, and repeat it on every subdomain you want excluded. Anthropic's help page says this signals that the site's future material should be left out of its training datasets. Use the same pattern with Claude-SearchBot or Claude-User to block those separately.
Does blocking ClaudeBot remove my site from Claude's answers?
Anthropic's help page ties ClaudeBot only to training data. Search visibility is tied to Claude-SearchBot and live fetches to Claude-User, and the page says disabling either of those may reduce your visibility. So a ClaudeBot block on its own is a training decision, and the answer-facing controls are the other two tokens.
Which robots.txt rules should I use for AI bots?
Write one group per token, because each provider splits its bots by job. For Anthropic that means three groups: ClaudeBot, Claude-SearchBot and Claude-User. A business that wants to be cited usually allows the search and user tokens and decides training separately.
Want to know how AI models currently describe your business?
We run a free visibility check across ChatGPT, Perplexity, Claude and Google AI Overviews, then show you exactly which signals are missing.
Book a visibility check