AI Crawlers

GPTBot: what it is, and how to allow or block it

By Nihanth Guntur · 2026-10-03

GPTBot is OpenAI's web crawler for training data: it fetches pages that may be used to train OpenAI's generative AI foundation models, and you control it with a GPTBot group in robots.txt. Blocking GPTBot is a training opt-out, not a search opt-out. ChatGPT search relies on a separate crawler, OAI-SearchBot, which needs its own rule.

What is GPTBot?

GPTBot is the crawler OpenAI uses to collect content for model training. It is one of several OpenAI agents, and each one answers to a different robots.txt token.

OpenAI's overview of its crawlers says GPTBot is used to make its generative AI foundation models more useful and safe, and that it crawls content that may be used in training those models. The same page says disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models.

Read that sentence carefully. It describes an opt-out from training. It does not say anything about whether your pages can appear in ChatGPT answers, because that is a different crawler's job.

Does blocking GPTBot remove you from ChatGPT search?

No. Blocking GPTBot does not block OAI-SearchBot, and OAI-SearchBot is the crawler OpenAI says surfaces websites in ChatGPT's search features. The two are controlled separately.

OpenAI states this directly: each setting is independent of the others, and a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training. That is the setup we recommend when a client wants to opt out of training: stay findable, decline training.

The reverse mistake is the costly one. OpenAI's page says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. A blanket rule written to stop training can therefore remove you from ChatGPT search too, if it also catches OAI-SearchBot.

There is a third agent. OpenAI says ChatGPT-User may visit a page when someone asks ChatGPT or a CustomGPT a question, that it is not used for crawling the web automatically, and that because these actions are initiated by a user, robots.txt rules may not apply.

TokenWhat OpenAI says it doesWhat a Disallow rule means
GPTBotCrawls content that may be used to train OpenAI's generative AI foundation modelsIndicates your content should not be used in training
OAI-SearchBotSurfaces websites in search results in ChatGPT's search featuresYour site will not be shown in ChatGPT search answers, though it can still appear as a navigational link
ChatGPT-UserMay visit a page when a user asks ChatGPT or a CustomGPT a question; not used for automatic crawlingOpenAI says robots.txt rules may not apply, because a user initiated the visit

What is the GPTBot user agent, and how does robots.txt match it?

In robots.txt you address GPTBot by its product token, written as User-agent: GPTBot. That token is what this article covers. For OpenAI's own description of each agent, go to its crawler page rather than a third-party list.

How the matching works is set by the Robots Exclusion Protocol, published as RFC 9309. Three of its rules matter for a GPTBot group:

  • Crawlers must use case-insensitive matching to find the group for their product token, so gptbot and GPTBot match the same group.
  • If no group matches, a crawler must obey the User-agent: * group, if there is one. With no GPTBot group, your general rules apply to GPTBot.
  • If more than one group matches, the matching groups' rules must be combined into one. Two scattered GPTBot sections are read as one.

Within a group, RFC 9309 says the most specific match must be used, meaning the rule with the most octets, and that when an allow rule and a disallow rule are equivalent, the allow rule should be used. That is how you can disallow a folder for GPTBot and still allow one page inside it.

How to block GPTBot without blocking ChatGPT search

Give each OpenAI token its own group, decide each one on purpose, and test the result. Here is the procedure we follow on client sites.

  1. Open your live robots.txt at the root of the domain. RFC 9309 says the file must be named /robots.txt, all lowercase, in the top-level path, UTF-8 encoded and served as text/plain.
  2. Search it for any group that names GPTBot, OAI-SearchBot or ChatGPT-User, and for a User-agent: * group with a broad Disallow. Note which group each OpenAI agent will actually fall into.
  3. Add an explicit group for each OpenAI token you have a view on, rather than relying on the wildcard group.
  4. Publish the file and confirm it returns status 200. RFC 9309 says that if robots.txt is unreachable because of server errors, in the 500-599 range, a crawler must assume complete disallow. A broken robots.txt blocks more than you intended.
  5. Allow time for the change. For search results, OpenAI says it can take about 24 hours from a robots.txt update for its systems to adjust.
  6. Check the firewall and CDN separately. robots.txt is a request; a bot rule at the edge is a block, and it can catch crawlers your robots.txt allows.

For the common goal of opting out of training while staying in ChatGPT search, the file looks like this:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

To allow GPTBot everywhere except a private section, keep the group and narrow the rule:

User-agent: GPTBot
Disallow: /members/
Allow: /

OpenAI also recommends allowing requests from its published IP ranges if you want to appear in search results. If a firewall sits in front of your site, that is the setting to look at.

Is robots.txt enough to block GPTBot?

robots.txt is a signal that a well-behaved crawler reads and obeys. It is not access control, and the standards say so plainly.

RFC 9309 says the rules are not a form of access authorization, and that the protocol is not a substitute for valid content security measures. Google's robots.txt guide makes the same point from the search side: the instructions in robots.txt cannot enforce crawler behaviour, and it is up to the crawler to obey them.

Google's guide adds a second limit worth knowing. A page that is disallowed in robots.txt can still be indexed if it is linked from other sites, and its URL can still appear in search results without a description. Google says that to keep a page out of its results you should use noindex, password-protect the page, or remove it.

So match the tool to the goal. To state a training preference to OpenAI, a GPTBot rule is the mechanism OpenAI documents. To keep content private from everyone, put it behind a login. robots.txt does neither job alone for a page that must stay private.

What we have seen when GPTBot rules meet a firewall

robots.txt is only half the check. A security rule at the edge can block an OpenAI crawler that robots.txt allows. We wrote up one case in how a firewall blocked AI crawlers.

On one SEO Autopilot engagement, a US healthcare analytics company on WordPress behind a firewall, the firewall returned 403 to OpenAI's and Anthropic's search crawlers on every page. The trigger was a generic bad-bot rule matching the substring "searchbot". Eleven other crawlers passed. The client's monitoring tool reported all sixteen AI crawlers as accessible, because it read robots rules instead of sending requests.

A robots.txt review alone would have reported that site as open to ChatGPT search. Only a real request showed the 403. If you are deciding on GPTBot, check OAI-SearchBot at the same time, at both layers.

How do I check whether GPTBot can reach my site?

Read your robots.txt groups for each OpenAI token, then send requests and look at the status codes your server returns. Both layers have to agree.

Our free AI visibility checker reads robots.txt rules for 11 AI crawlers and requests your homepage once with each crawler's user agent. That request comes from the checker, not from OpenAI, and it covers the homepage only, so treat it as a first pass. It can surface a robots rule or an edge block worth investigating; it cannot prove that GPTBot or OAI-SearchBot will reach every page.

For the wider picture of who else is crawling and which tokens to list, see our AI crawlers list. If you want this checked across a whole site every month, alongside the content work that earns citations, that is part of SEO Autopilot.

Frequently Asked Questions

What is GPTBot used for?

OpenAI says GPTBot crawls content that may be used in training its generative AI foundation models. Disallowing it in robots.txt indicates your content should not be used for that training.

Does blocking GPTBot stop my site appearing in ChatGPT?

Not on its own. ChatGPT search uses OAI-SearchBot, which OpenAI controls with a separate robots.txt setting. You can disallow GPTBot and still allow OAI-SearchBot.

How do I block GPTBot in robots.txt?

Add a group that starts with User-agent: GPTBot followed by Disallow: / to block the whole site, or Disallow with a specific path to block one section. Give OAI-SearchBot its own group if you want to stay in ChatGPT search.

Is the GPTBot user agent case-sensitive in robots.txt?

No. RFC 9309 says crawlers must use case-insensitive matching to find the group for their product token, so GPTBot and gptbot match the same group.

How long does a robots.txt change take to affect ChatGPT search?

For search results, OpenAI says it can take about 24 hours from a robots.txt update for its systems to adjust.

Free AI Visibility Check

Is OpenAI reaching your site, or just your robots.txt?

Run the free checker to see your robots.txt rules for 11 AI crawlers and how your homepage answers each crawler's user agent.