AI Crawler User Agents: ChatGPT, Claude, Gemini and Perplexity
Most robots.txt files written for AI are wrong in the same way: they block the training crawler and leave the search crawler alone, then the site owner wonders why nothing changed.
Those are different bots with different jobs. One decides whether your pages help train a future model. The other decides whether you can appear in an answer today. Blocking the first costs you nothing in visibility. Blocking the second removes you from the answer entirely.
Here is every agent the four engines use, what each one does, and which ones actually matter.
The distinction that matters
Training crawlers
GPTBot, ClaudeBot, Google-Extended. They collect content that may be used to train future models. Blocking them is a licensing and policy decision. It does not remove you from answers given today.
Search and retrieval crawlers
OAI-SearchBot, Claude-SearchBot, PerplexityBot. They decide whether your pages can be surfaced and cited in answers right now. Blocking these is what makes a brand disappear.
There is a third category: agents that fetch a page because a user asked a question that needs it. Those behave differently again, and two of them say plainly that they may not follow robots.txt.
If the terms here are unfamiliar, start with what crawlability means.
OpenAI — ChatGPT
| Token | What it does |
|---|---|
OAI-SearchBot | Surfaces sites in ChatGPT's search features. This is the one that governs whether you can appear. |
GPTBot | Crawls content that may train OpenAI's foundation models. |
ChatGPT-User | Handles user-initiated actions — visiting a page in response to a question, and GPT Actions. Not automatic crawling, and robots.txt rules may not apply. |
OAI-AdsBot | Checks the safety of landing pages submitted as ChatGPT ads. Visits only submitted ad pages; its data is not used for training. |
Two things worth knowing. ChatGPT-User does not determine search
eligibility — OAI-SearchBot does, so blocking the wrong one achieves nothing.
And opting out of OAI-SearchBot doesn't erase you completely: sites can still
appear as plain navigational links.
OpenAI publishes IP ranges for each agent as JSON, so you can verify a request really came from them rather than something wearing the user-agent string.
Anthropic — Claude
| Token | What it does |
|---|---|
ClaudeBot | Collects content that may help train Anthropic's models. Blocking it signals your material should be excluded from training sets. |
Claude-SearchBot | Navigates the web to improve search result quality. Blocking it stops indexing for search, which reduces visibility and accuracy in responses. |
Claude-User | Accesses sites when a Claude user asks a question that needs them. Blocking it stops retrieval for those requests. |
The pattern is the same as OpenAI's: one for training, one for the search index, one for live user requests.
Perplexity
| Token | What it does |
|---|---|
PerplexityBot | Surfaces and links sites in Perplexity results. Explicitly not used to crawl content for AI foundation models. |
Perplexity-User | Visits a page when a user asks a question, so the answer can be accurate and link to you. Because a user triggers it, it generally ignores robots.txt. |
Perplexity is the cleanest of the four: no training crawler at all. If you block
PerplexityBot you are only removing yourself from results, with nothing gained
on the training side.
Google — Gemini
| Token | What it does |
|---|---|
Google-Extended | A robots.txt control token. Not a crawler. |
This is the one that catches people out.
What it controls: whether content Google has already crawled can be used to train future Gemini models, and whether it can be used for grounding — where content from the Search index is handed to the model at prompt time in Gemini Apps and in Grounding with Google Search on Vertex AI.
What it does not control: your inclusion in Google Search. Google states plainly that the token is not a ranking signal and does not affect whether you appear in ordinary search results.
So Google-Extended is the only lever here that is purely a policy choice, with
no crawling behavior attached to it.
A robots.txt that does what most brands want
If the goal is to be found and cited, allow everything that retrieves and leave the training decision to policy:
# Search and retrieval — allow these if you want to appear in answers
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Training — a licensing decision, not a visibility one
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
If you want to stay out of training sets while remaining fully visible in
answers, change the three training entries to Disallow: / and leave the search
group alone. That combination is coherent: it says read me and cite me, but
don't learn from me.
What doesn't work is the reverse — blocking the search crawlers and keeping the training ones. That removes you from today's answers in exchange for a benefit that arrives, at best, in some future model.
On the user-triggered agents
ChatGPT-User, Claude-User and Perplexity-User fetch pages because a person
asked a question that needed them. Two of the three vendors say outright that
robots.txt may not apply to those requests, on the reasoning that a human
initiated it rather than a crawler.
The practical consequence: you cannot reliably keep your pages out of a user's hands through robots.txt. If that matters, it's an access-control problem, not a robots.txt one.
Verify before you trust the user agent
A user-agent string is just a header. Anything can claim to be GPTBot.
OpenAI publishes IP ranges per agent as JSON files you can check a request against. If you are making decisions based on crawler traffic — rate limiting, serving different content, reporting crawl activity to a client — verify by IP rather than by the string. The string is a claim; the IP range is evidence.
This page will go out of date
Agents get added, renamed and split. OAI-SearchBot and Claude-SearchBot both
arrived after the original training crawlers, and the user-triggered agents
arrived after that.
Every entry here was checked against each vendor's own documentation on 13 October 2026. If you are reading this much later, check the primary sources before acting on it — and treat any robots.txt you wrote more than six months ago as something to revisit rather than something that's still correct.
One more thing this list does not cover: crawlers belonging to engines outside these four, and dataset crawlers that feed many models at once. They exist and they are worth knowing about. But these ten are the ones that decide whether ChatGPT, Claude, Gemini and Perplexity can see you.