PerplexityBot and Perplexity-User: what each one fetches, and what the stealth crawling report means for your robots.txt
Perplexity documents two user agents. PerplexityBot indexes pages for Perplexity search and obeys robots.txt. Perplexity-User fetches a page when someone asks a question, and Perplexity's own docs say it generally ignores robots.txt. In August 2025 Cloudflare reported an undeclared crawler it tied to Perplexity, and Perplexity denied it. The two user-agent strings, what each one fetches, what the report claimed, and the robots.txt and firewall rules a business site should ship.
Short answer
PerplexityBot indexes pages for Perplexity's search results and follows robots.txt. Perplexity-User fetches a page live when someone asks a question, and Perplexity's docs say it generally ignores robots.txt. For a business that wants to be cited, allow both. To keep Perplexity out, you need a firewall rule, not a robots.txt line.
Perplexity documents two user agents, and by its own account only one of them reads your robots.txt.
The Perplexity Crawlers page, fetched 9 October 2026, says of the second one: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."
That line is the reason people type "perplexitybot robots.txt" into Google and come away confused.
Then there is the report. On 4 August 2025 Cloudflare published an account of a third, undeclared crawler it tied to Perplexity. Perplexity denied it.
Here is how the pieces fit, and what to put in the file.
1. The two names, in Perplexity's words
PerplexityBot builds the index. The documentation says it exists to surface and link websites in Perplexity's search results. It also says PerplexityBot is not used to crawl content for AI foundation models.
The full string, per the same page:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity-User fetches a page while someone waits. When a person asks Perplexity a question, it may visit your page to answer and link to it. Perplexity says this agent is not used for crawling or for training either.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
Each has its own IP list. The PerplexityBot file, fetched 9 October 2026, carries a creation time of 7 February 2025. The Perplexity-User file carries 17 October 2025.
So by its own timestamp, the current Perplexity-User IP file postdates the Cloudflare report. Check the files, not a blog post. This one included.
2. An indexer and a fetcher are different machines
This distinction decides everything else.
An indexer crawls on its own schedule. It reads robots.txt, decides which of your pages to keep, and stores them for later answers. PerplexityBot is that.
A fetcher only moves because a person asked something. It loads one page, now, to answer one question. Perplexity-User is that.
Perplexity is not alone in treating the second kind as outside robots.txt. Google's user-triggered fetchers page, last updated 19 August 2026 and read on 9 October 2026, says "these fetchers generally ignore robots.txt rules." OpenAI's bots page, read the same day, says robots.txt rules may not apply to ChatGPT-User because a user started the action.
The logic across all three vendors is the same. A fetcher is treated as a browser with a person behind it, and robots.txt was written for crawlers.
You can disagree with that logic. Plenty of publishers do. But it is the documented position of all three vendors above, so plan around it.
3. What Cloudflare said it saw
Here is the complication.
Cloudflare's report, dated 4 August 2025 and read on 9 October 2026, starts with customer complaints. Sites had disallowed Perplexity in robots.txt and blocked both declared agents with firewall rules. Perplexity still seemed to know what was on them.
So Cloudflare bought brand-new domains, never indexed or linked anywhere, and gave each a robots.txt that disallowed every bot. Then it asked Perplexity about them.
Cloudflare says Perplexity answered with details of the restricted content.
It then describes what it saw at the network level:
- the declared Perplexity-User agent at 20 to 25 million requests a day, by Cloudflare's count
- an undeclared agent presenting as Chrome 124 on a Mac at 3 to 6 million requests a day, by the same count
- that undeclared agent rotating through IPs outside Perplexity's published ranges, and across different networks
- robots.txt ignored, or sometimes not fetched at all
Cloudflare's conclusion: "we have de-listed them as a verified bot." It also added signatures for the stealth crawler to its managed rule that blocks AI crawling, which it said was available to free plans too.
It ran the same test on ChatGPT for contrast. Cloudflare says ChatGPT-User fetched robots.txt, saw the disallow, and stopped.
4. What Perplexity said back
Perplexity also answered on its own blog. This post does not paraphrase that response, because we did not read it in full.
What we could read is TechCrunch's report from 4 August 2025. In it, Perplexity spokesperson Jesse Dwyer dismissed Cloudflare's post as a "sales pitch." He said the screenshots showed no content had been accessed, and in a follow-up claimed the bot Cloudflare named was not Perplexity's.
That is the record. One company published network data. The other disputed what the data showed and who owned the traffic.
NovaTechRay cannot settle it. We did not run the test, and we do not have Cloudflare's logs.
What we can say is that the dispute matters less to a business site than the headline suggests. Section 5 is why.
5. What robots.txt can and cannot do here
robots.txt is a request. The standard says so in plain words: RFC 9309, read on 9 October 2026, states that "These rules are not a form of access authorization."
So the file does three jobs at most:
- it tells a declared indexer what not to store
- it tells a declared indexer how you feel about the rest
- it gives you a written preference to point at later
It does not stop a fetcher that the vendor says ignores the file. It does not stop an agent that does not declare itself. If Cloudflare's account is right, both of those were in play.
That makes the real question simple. Do you want to be in Perplexity's answers, or out of them?
6. The file we would ship
Allow both.
For a clinic, a restaurant, an agency or a hotel, Perplexity reading your pages is the goal. Its answers cite sources. A Perplexity-User fetch is a customer asking about you, live, with your link in the reply.
So the file is the same one our AI crawlers and robots.txt guide recommends: one wildcard group, no named Perplexity group at all.
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml
Then check your edge. A managed AI-crawler rule at your CDN can block Perplexity before your robots.txt is ever read. Our Cloudflare decision tree walks through which switches touch a business site.
If you are a publisher and want Perplexity out, the file alone will not do it:
User-agent: PerplexityBot
Disallow: /
That handles the indexer. Perplexity says changes can take up to 24 hours to reflect.
For Perplexity-User you need a firewall rule on the token plus its IP file. For anything undeclared, you are relying on your bot management vendor, which is what Cloudflare's report was selling, fairly or not.
7. Check your own logs before you decide
Do not decide from a headline. Decide from your access logs.
- Filter for both tokens. Search for PerplexityBot and Perplexity-User separately. They answer different questions.
- Match the IPs. Compare each hit against the matching JSON file. A real token from an unlisted IP is either a spoof or something Perplexity did not declare. Treat both as unverified.
- Look for robots.txt requests. A declared indexer should fetch it. Note how often.
- Ask Perplexity about yourself. Run five questions a customer would ask. See whether you are cited, and which page.
We check every page a crawler can reach the same way across the VALORAE group: the group site explains how a new page gets found, from blog index to sitemap.
NovaTechRay will update this post if Perplexity's documentation changes, or if a dated, readable response from Perplexity becomes available to us.
The same breakdown for OpenAI and Anthropic is in GPTBot, OAI-SearchBot and ChatGPT-User and ClaudeBot, Claude-User and Claude-SearchBot.
The short version
- Know that PerplexityBot indexes and Perplexity-User fetches for a person.
- Read Perplexity's docs: the user fetcher generally ignores robots.txt.
- Allow both if you want Perplexity to cite you.
- Block PerplexityBot in robots.txt only if you want out of its index.
- Use a firewall rule on token and IP for anything robots.txt cannot stop.
- Verify every hit against Perplexity's published IP files.
- Check your CDN's AI-crawler setting before blaming the file.
This is the work NovaTechRay does at novatechray.com.
Frequently asked questions
What is PerplexityBot?
PerplexityBot is Perplexity's crawler for its search index. Perplexity's crawler documentation, read on 9 October 2026, says it is designed to surface and link websites in Perplexity's search results and is not used to crawl content for AI foundation models. Perplexity recommends allowing it in robots.txt if you want to appear in its results.
What is the PerplexityBot user agent?
Per Perplexity's crawler documentation, read on 9 October 2026, the full string is Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot). Its published IP ranges are at perplexity.com/perplexitybot.json. Match on both the token and the IP list, because anyone can copy a user-agent string.
How do I block PerplexityBot in robots.txt?
Add a group for it: User-agent: PerplexityBot, then Disallow: /. Perplexity's documentation says changes can take up to 24 hours to reflect. That group does not stop Perplexity-User, which Perplexity says generally ignores robots.txt because a person requested the fetch, so a full block needs a firewall rule on the Perplexity-User token and its published IP ranges as well.
Does Perplexity-User follow robots.txt?
Generally no, by Perplexity's own account. Its crawler documentation, read on 9 October 2026, says that since a user requested the fetch, Perplexity-User generally ignores robots.txt rules. Google's user-triggered fetchers and OpenAI's ChatGPT-User are documented the same way, so this is how the category works across vendors, not a Perplexity exception.
Where is the PerplexityBot user agent documentation?
Perplexity publishes it at docs.perplexity.ai/guides/bots, titled Perplexity Crawlers. As read on 9 October 2026, the page lists both user agents with their full strings, links the two IP range files, and gives firewall setup steps for Cloudflare WAF and AWS WAF.
Is Perplexity's crawler using stealth user agents?
Cloudflare said so on 4 August 2025. It reported an undeclared crawler presenting as Chrome on a Mac, rotating IPs and networks after Perplexity's declared agents were blocked, and de-listed Perplexity as a verified bot. Perplexity's spokesperson told TechCrunch the same day the post was a sales pitch and the bot named was not Perplexity's.
Want to know how AI models currently describe your business?
We run a free visibility check across ChatGPT, Perplexity, Claude and Google AI Overviews, then show you exactly which signals are missing.
Book a visibility check