AI crawlers are the bots AI companies send to read your pages, and they come in three kinds: training crawlers such as GPTBot and ClaudeBot, search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, and user-triggered fetchers such as ChatGPT-User. Each vendor documents its own robots.txt tokens, so you can allow search while opting out of training.
What are AI crawlers, and why are there three kinds?
An AI crawler is any automated agent an AI company uses to fetch web pages. The useful split is by job, because each job affects your visibility differently.
Anthropic's help centre puts it plainly: Anthropic uses a variety of robots to gather data from the public web for model development, to search the web, and to retrieve web content at users' direction. That sentence is the whole taxonomy. A training crawler collects content that may end up in a model. A search crawler indexes pages so an AI search product can find and cite them. A user-triggered fetcher visits one page because a person asked a question in a chat.
OpenAI draws the same lines. Its crawler overview says each setting is independent of the others, so a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training. That independence is the reason a single "block all AI" rule is a blunt tool: it removes you from AI search answers at the same time as it opts you out of training.
The AI crawler list: user agents from each vendor's own documentation
| robots.txt token | Vendor | Kind | What the vendor says it does | What blocking or allowing it means |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | Crawls content that may be used in training OpenAI's generative AI foundation models | Disallowing it indicates content should not be used in training |
| OAI-SearchBot | OpenAI | Search | Surfaces websites in search results in ChatGPT's search features | OpenAI recommends allowing it; changes can take about 24 hours |
| ChatGPT-User | OpenAI | User-triggered | May visit a page when a user asks ChatGPT or a CustomGPT a question; not used for automatic crawling | OpenAI says robots.txt rules may not apply |
| ClaudeBot | Anthropic | Training | Collects web content that could potentially contribute to model training | Restricting it signals future materials should be excluded from training |
| Claude-SearchBot | Anthropic | Search | Navigates the web to improve search result quality for users | Disabling it may reduce visibility in user search results |
| Claude-User | Anthropic | User-triggered | Accesses websites when individuals ask Claude questions | Disabling it prevents retrieval in response to a user query |
| PerplexityBot | Perplexity | Search | Surfaces and links websites in Perplexity search results; not used to crawl content for AI foundation models | Perplexity recommends allowing it; changes can take up to 24 hours |
| Perplexity-User | Perplexity | User-triggered | Supports user actions within Perplexity; not used for web crawling or training | Perplexity says it generally ignores robots.txt rules |
| Google-Extended | Control token | No separate HTTP user agent string; crawling uses existing Google user agents | Does not affect inclusion in Google Search or ranking |
Note the user-triggered row. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a person asked for the page. Anthropic's page instead says Claude-User lets site owners control which sites can be reached this way. And only the training rows are about training: blocking a search crawler is a visibility decision.
Names that are not in this table, such as Bytespider, are not covered by any page linked here, so we do not describe them. Before writing a rule for any crawler, read its owner's own documentation.
What does ClaudeBot do, and how is it different from Claude-SearchBot?
ClaudeBot is Anthropic's training crawler; Claude-SearchBot is its search crawler. Blocking the first is an opt-out from training. Blocking the second can cost you visibility in Claude's search results.
Anthropic's help centre says ClaudeBot collects web content that could potentially contribute to model training, and that restricting ClaudeBot signals that a site's future materials should be excluded from training datasets.
Claude-SearchBot is described differently. It analyses online content specifically to improve the relevance and accuracy of search responses, and Anthropic says disabling it prevents its system from indexing your content for search optimisation, which may reduce your visibility in user search results.
Anthropic says its bots honour industry standard directives in robots.txt, and that you should add rules for every subdomain you want to opt out.
Where do PerplexityBot and Google-Extended fit?
PerplexityBot is a search crawler, and Google-Extended is not a crawler at all but a robots.txt control token. Treat them as two different decisions.
Perplexity’s crawler documentation says PerplexityBot is designed to surface and link websites in search results on Perplexity, and that it is not used to crawl content for AI foundation models. Perplexity recommends allowing PerplexityBot in robots.txt and permitting requests from its published IP ranges if you want to appear in its results. If you run a web application firewall, the same page says you may need to explicitly allow Perplexity's bots.
Google-Extended is the odd one out. Google’s common crawlers page says it has no separate HTTP request user agent string; crawling is done with existing Google user agent strings, and the robots.txt token is used in a control capacity. The same page says Google-Extended does not impact a site's inclusion in Google Search, nor is it used as a ranking signal. So you will not see a separate Google-Extended user agent in your logs, and disallowing it does not remove you from Google Search. Crawling preferences addressed to Googlebot are what affect Google Search, including Discover and all Search features.
Google's page also warns that user agent strings can be spoofed, and that log searches should use wildcards for the version number.
How do you allow AI crawlers in robots.txt?
Name each AI crawler you want by its robots.txt token, give it the access you intend, and keep training and search decisions separate. Here is the procedure we follow on client sites:
- Decide your policy per kind, not per company: training (GPTBot, ClaudeBot), search (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and user-triggered fetchers. Treat the Google-Extended token as a separate decision; Google's page says it does not affect Search inclusion or ranking.
- Open robots.txt at the root of every subdomain you care about, since Anthropic's page says rules are needed for each subdomain you want to opt out.
- Add a User-agent group for each token with the Allow or Disallow rules you chose. Do not rely on a wildcard group to express an AI policy you care about.
- Check for older rules that contradict the new ones, such as a blanket Disallow left from a staging site.
- Wait before judging the result. OpenAI says search changes can take about 24 hours to adjust, and Perplexity says up to 24 hours.
- Test that requests actually reach the site, then read your server or firewall logs for the real crawlers' responses.
A site that wants to be found in AI search but opt out of training might use a set of groups like the one below. It is an illustration of the structure, not a recommendation for every site:
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
Why is robots.txt not the whole answer for AI crawlers?
Because robots.txt states permission, and a firewall decides whether the request is served. The two can disagree, and only one of them shows up in a robots.txt report.
Cloudflare’s verified bots documentation describes a verified bot as one Cloudflare has confirmed is transparent about who it is and does not abuse that access, identified through a Web Bot Auth signature, a published IP list with a stable user agent, or reverse DNS. It says that under the taxonomy introduced on July 1, 2026, there is no longer a meaningful distinction between "AI Search" and traditional search, and that customers now have the option to configure AI bot policies to define what they block and allow. If your site sits behind Cloudflare, those settings are a second place where your AI crawler policy lives, and they need to agree with robots.txt.
What we have seen: on one SEO Autopilot engagement, a US healthcare analytics company on WordPress behind a firewall, the firewall returned 403 to OpenAI's and Anthropic's search crawlers on every page. The trigger was a generic bad-bot rule matching the substring "searchbot". Eleven other crawlers passed. The client's monitoring tool reported all sixteen AI crawlers as accessible, because it read robots rules instead of sending requests. We wrote up how that happened in AI crawlers blocked by a firewall.
Our free AI Visibility Checker reads robots.txt rules for 11 AI crawlers and requests your homepage once with each crawler's user agent. Be clear about the limits. That request is the checker's own, sent from Cloudflare, not the real crawler's, and it covers the homepage only. A firewall that verifies bots by IP or signature may treat it differently from the real crawler, so it cannot prove a crawler can reach your site. Your server logs are the proof.
Access is only the first step. Once crawlers can read your pages, what they find there decides whether you are cited, which we cover in where ChatGPT gets its information. Allowing a crawler does not promise a citation; it removes one reason you could not get one. Our AI SEO services page explains how we work on AI search visibility.
Frequently Asked Questions
Which AI crawlers should I allow?
If you want to appear in AI search answers, allow the search crawlers: OAI-SearchBot, Claude-SearchBot and PerplexityBot. Whether you allow training crawlers such as GPTBot and ClaudeBot is a separate decision, and OpenAI documents them as independent settings.
Does blocking GPTBot remove my site from ChatGPT search?
Not according to OpenAI's crawler overview. It describes GPTBot as the training crawler and OAI-SearchBot as the crawler that surfaces websites in ChatGPT search, and says each setting is independent. Sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
Does Google-Extended affect my Google rankings?
Google's common crawlers page says Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal. It has no separate user agent string; it is a robots.txt token used in a control capacity.
Do user-triggered AI fetchers obey robots.txt?
The vendors differ. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt rules. Anthropic says Claude-User lets site owners control which sites can be accessed through user-initiated requests.
How long do robots.txt changes take to reach AI crawlers?
OpenAI says it can take about 24 hours from a robots.txt update for its search systems to adjust, and Perplexity says it may take up to 24 hours for its systems to reflect changes.
Is blocking AI crawlers by IP address enough?
Anthropic's help centre says blocking the IP addresses its bots operate from may not work correctly or persistently guarantee an opt-out, because it impedes their ability to read your robots.txt file. For an opt-out, use robots.txt.