GPTBot and OAI-SearchBot have different purposes. GPTBot crawls content that may be used to train OpenAI's generative AI foundation models. OAI-SearchBot helps surface websites in ChatGPT search. If your business wants to be discoverable through search without allowing GPTBot's training-related crawling, OpenAI provides separate robots.txt controls for that choice.
A blanket instruction to block every AI bot can therefore affect more than training preferences. Equally, allowing every crawler is not a requirement for earning recommendations. Start by deciding which uses of your public content you want to permit, then verify that your website and security settings implement that policy.
This guide follows OpenAI's official crawler documentation, reviewed on 9 October 2026. Crawler details and behaviour can change, so check the live documentation before altering production access rules.
GPTBot, OAI-SearchBot and ChatGPT-User Explained
GPTBot: Training-related crawling. — OpenAI describes GPTBot as a crawler for content that may be used in training its generative AI foundation models. Disallowing GPTBot indicates that your site's content should not be used for that purpose.
OAI-SearchBot: Search discovery. — This crawler is used to surface websites in ChatGPT's search features. Its permission is independent of the GPTBot setting.
ChatGPT-User: User-initiated visits. — This agent may visit a page when a user asks ChatGPT or a Custom GPT to do something involving it. It is not an automatic web crawler and is not the control for search inclusion.
OpenAI states that sites opting out of OAI-SearchBot will not be shown in ChatGPT search answers, although they can still appear as navigational links. Do not turn that into a claim that blocking the crawler erases every possible mention of a business: third-party sources, user-provided information and non-search responses are separate considerations.
For ChatGPT-User, OpenAI says robots.txt rules may not apply because the actions are initiated by users. Use OAI-SearchBot to manage automatic search crawling. Use proper authentication and access controls to protect private material; do not rely on a robots.txt entry to keep it confidential.
Can You Block GPTBot and Still Allow ChatGPT Search?
Yes. OpenAI explicitly describes allowing OAI-SearchBot while disallowing GPTBot. Search access and training-related crawling are not the same permission. Enabling GPTBot is not a documented prerequisite for ChatGPT search inclusion.
The example below permits search crawling of public pages while excluding two illustrative sensitive routes and disallowing GPTBot. Replace the example paths with the appropriate routes for your site, and merge the groups into your existing robots.txt rather than replacing the entire file.
User-agent: OAI-SearchBot
Allow: /
Disallow: /account/
Disallow: /checkout/
User-agent: GPTBot
Disallow: /Robots rules guide compliant crawlers; they are not a security boundary. If an account page contains private information, it needs authentication whether or not that route appears in robots.txt. Also note that a specific user-agent group does not simply inherit restrictions from your wildcard group. Review any existing exclusions when adding a new group.
Disallowing GPTBot is a forward-looking crawling preference, not proof that previously collected material has been deleted or removed from an already trained model. Do not make retrospective deletion claims based solely on a file change.
When Would You Allow or Disallow Both?
A business that permits both search crawling and potential training use can allow both agents, subject to appropriate route restrictions and its content policy. A publisher or regulated organisation may instead decide to exclude search crawling as well as training-related crawling. Make that choice with the people responsible for the content and its permitted uses.
To disallow both agents, use separate groups such as the following. The trade-off is intentional: you are opting out of OpenAI search crawling as well as GPTBot crawling, not just restricting training.
User-agent: OAI-SearchBot
Disallow: /
User-agent: GPTBot
Disallow: /A more targeted policy can permit public service pages while excluding selected paths. Review the actual groups, path rules and host serving them before making changes. Keep your sitemap declaration and rules for other search engines intact unless you also intend to alter their access.
Check More Than the robots.txt File
A correct robots.txt file does not help if your CDN, firewall or hosting platform returns a challenge page to the crawler. A browser visit by your team may succeed while an automated request receives a 403, a rate-limit response or HTML asking it to run JavaScript.
For important public pages, check the response status, content type, redirects and response body. A 200 status alone is not enough: it may describe an empty application shell, a challenge page or an error message. Confirm that the answer-bearing content is available in a form the requesting crawler can retrieve.
Pay attention to hostname differences. A rule served on one host does not necessarily apply to another, and redirects may lead a crawler through several hosts. Test the canonical host that serves your public content.
If you run WordPress or another CMS, a plugin may generate robots.txt dynamically. Security plugins and bot-management services can also change crawler behaviour independently of that file. Treat the deployed response as the evidence, rather than assuming the configuration screen tells the whole story.
Verify a Crawler Before Allowlisting It
A user-agent string can be copied by any HTTP client. Seeing OAI-SearchBot or GPTBot in a log entry does not, by itself, prove that the request came from OpenAI.
OpenAI publishes separate address lists for OAI-SearchBot, GPTBot and ChatGPT-User. Use the current published ranges, or an appropriately maintained verified-bot facility from your security provider, when validating crawler identity.
If your site sits behind a proxy, use the client address determined by your trusted proxy configuration. Do not trust arbitrary forwarded headers from the public internet. Have your developer make any allowlist rule as narrow as needed, rather than disabling security checks for every request claiming to be a bot.
A Practical Access-Testing Workflow
Step 1: Record the current policy. — Save the deployed robots.txt response and identify the exact user-agent groups and route exclusions. Agree whether the goal is search access, training-related access or both.
Step 2: Test a priority page. — Request a real service or article URL and inspect its status, redirects and returned HTML. Compare ordinary requests with requests carrying the relevant user-agent string.
Step 3: Inspect genuine crawler requests. — Check logs for requests from the published address ranges. Look for challenges, blocked paths, unexpected redirects and rate limits.
Step 4: Make a targeted correction. — Change the specific robots group, route restriction or verified-crawler security rule responsible for the problem. Preserve unrelated restrictions.
Step 5: Retest and document. — Check the deployed response again, record the change time and watch subsequent crawler requests. Then monitor relevant search prompts separately.
A developer can use a command like this to inspect a response body and its headers. Replace example.com with your own host and choose a public priority page. The command stores the body locally so it can be inspected.
curl -sS -D - -o /tmp/ai-page.html \
-A 'OAI-SearchBot' \
https://example.com/your-priority-page/This is a user-agent simulation, not a genuine OpenAI request. Curl does not automatically enforce robots.txt, and a request from your machine cannot validate an IP-based rule for OpenAI's infrastructure. Use it to diagnose the response, then check real crawler evidence or your security provider's verified-bot tooling.
OpenAI says its search systems may take approximately 24 hours to adjust after a robots.txt update. That is not a promise of a citation within 24 hours. Crawling, retrieval and answer selection remain separate steps.
Access Is Necessary to Check, Not Enough to Win Citations
Allowing OAI-SearchBot removes one potential search-access obstacle. It does not guarantee that a page will be selected, cited or recommended. The page still needs to be relevant to the buyer's question and provide useful, trustworthy information.
Do not confuse a crawler visit with model training, or a search crawl with a successful recommendation. Likewise, an llms.txt file does not override robots.txt, bypass authentication or establish that ChatGPT will cite your site. Crawl evidence tells you about access; answer observations tell you about visibility.
Once access is working, assess your service descriptions, factual evidence, author expertise and the third-party sources that explain your business. Our guide to getting cited by ChatGPT covers the wider content and authority work.
Use a stable set of buyer questions and repeat checks rather than relying on one screenshot. Our AI visibility tools guide for UK SMEs explains how to choose monitoring that preserves prompts, dates and citations. For a broader diagnosis, see the AI SEO audit guide.
The Decision for Your Business
If you want public pages discoverable in ChatGPT search but do not want GPTBot's training-related crawling, allow OAI-SearchBot and disallow GPTBot using the separate groups. If your policy permits or excludes both uses, implement that choice explicitly and test the deployed behaviour.
The important sequence is policy, verified access, useful content and repeated measurement. None of those steps can be replaced by a blanket promise that allowing AI bots will make your business rank.
Frequently Asked Questions
Q: Does blocking GPTBot stop my business appearing in ChatGPT search?
A: Not by itself. OpenAI documents independent settings for GPTBot and OAI-SearchBot. You can disallow GPTBot while allowing the search crawler, although search access does not guarantee inclusion or citations.
Q: Is ChatGPT-User the same as OAI-SearchBot?
A: No. ChatGPT-User handles certain user-initiated visits, while OAI-SearchBot manages automatic search crawling. OpenAI says robots.txt rules may not apply to the user-initiated actions.
Q: Can robots.txt protect confidential pages?
A: No. Robots.txt is not access control. Protect confidential content with authentication and appropriate server-side permissions, irrespective of crawler preferences.
Unsure whether crawler restrictions are limiting your AI search visibility? Panovista Marketing can help you inspect access evidence and prioritise the technical and content changes that matter.
