MonitorMyGEO uses essential storage for account and security functions. Google Analytics and Microsoft Clarity load only if you allow analytics. Read our Cookie Policy.
How Does ChatGPT Find Websites? OAI-SearchBot, Search Retrieval and Citations Explained
A practical guide to how websites become eligible for ChatGPT search, what OAI-SearchBot does, how crawler access differs from ranking, and what publishers should verify.
Short answer: A public website can appear in ChatGPT search, but discovery starts with technical eligibility rather than a special “GEO submission” button. OpenAI says publishers that want their content included in ChatGPT search summaries and snippets should not block OAI-SearchBot, and should also make sure their host or CDN allows traffic from OpenAI’s published searchbot IP ranges. That makes content eligible to be discovered; it does not guarantee placement, citation, or recommendation.
If you are new to the subject, start with our complete guide to Generative Engine Optimization. This article focuses on one narrower question: what has to happen between publishing a web page and that page becoming usable in a ChatGPT search experience?
OAI-SearchBot is the crawler for ChatGPT search discovery
OpenAI’s current publisher guidance separates search discovery from model training. OAI-SearchBot is the crawler publishers should allow when they want public content to be available for ChatGPT search experiences. GPTBot, by contrast, is the user agent publishers can disallow when they want to opt pages out of potential model training.
That distinction matters. Blocking GPTBot is not the same instruction as blocking OAI-SearchBot. A publisher may have different preferences for search visibility and model training, so robots.txt policies should be written deliberately rather than treating every OpenAI user agent as one thing.
OpenAI also says any public website can appear in ChatGPT search. For content to be included in summaries and snippets, however, the publisher should make sure OAI-SearchBot is not blocked. OpenAI separately advises checking whether the site host or content delivery network allows traffic from its published searchbot IP addresses.
A practical model: crawlability → retrieval → answer → citation
OpenAI does not publish a complete ranking algorithm for ChatGPT search. So it would be misleading to describe a precise internal pipeline as documented fact. A useful operational model for publishers is instead to separate four observable stages.
1. Crawlability and access
The crawler first needs to be able to reach the relevant public page. Robots.txt is one layer, but it is not the only one. A CDN, web application firewall, bot-protection service, authentication wall, rate limiter, geo rule or JavaScript challenge can also interfere with automated access.
OpenAI’s crawler troubleshooting guidance explicitly calls out robots.txt, web protection and bot mitigation. It also publishes a stable searchbot IP list for infrastructure that needs IP-based allowlisting.
This creates a simple diagnostic rule: “allowed in robots.txt” does not necessarily mean “successfully accessible.” A site can invite a crawler at the robots layer and then have Cloudflare or another security layer reject the request with a 403.
Frequently asked questions
What is OAI-SearchBot?
OAI-SearchBot is the OpenAI crawler publishers should allow when they want public content to be eligible for ChatGPT search summaries and snippets.
Is OAI-SearchBot the same as GPTBot?
No. OpenAI documents OAI-SearchBot for search discovery and GPTBot separately for potential model-training use, allowing publishers to make different access choices.
Does allowing OAI-SearchBot guarantee a ChatGPT citation?
No. OpenAI says search placement is not guaranteed. Allowing the crawler addresses access and eligibility, not guaranteed ranking, citation or recommendation.
Can a CDN block OAI-SearchBot even when robots.txt allows it?
Yes. OpenAI advises checking web protection, bot mitigation, firewalls, CDNs and rate limits in addition to robots.txt.
How can I identify traffic from ChatGPT search?
OpenAI says ChatGPT search referral URLs automatically include the UTM parameter utm_source=chatgpt.com, which can be tracked in analytics.
Research and editorial team at MonitorMyGEO covering Generative Engine Optimization, AI visibility, citations, measurement methodology and AI Readiness.
Monitor ChatGPT visibility changes by versioning the prompt panel, keeping raw answer evidence and separating real comparison periods from changes to the measurement setup.
Track ChatGPT citations by storing returned source URLs with the exact prompt and answer, normalizing domains and comparing source patterns across repeated runs.
There is no public formula that lets a marketer predict which brands ChatGPT will recommend. What can be measured are eligibility, retrieved evidence, answer context and repeated recommendation patterns.
Track ChatGPT brand mentions with a controlled prompt panel, clear mention rules and retained answer evidence. The method is more reliable than manually checking a few conversations.
If you want to inspect the wider technical layer rather than only OpenAI access, use the AI Visibility Audit Checklist as a companion review.
2. Search retrieval and relevance
Being crawlable is eligibility, not a ranking strategy. OpenAI states that ChatGPT search ranks results using multiple factors intended to help users find relevant and reliable information, and explicitly says placement is not guaranteed.
That is an important boundary for GEO work. Adding an Allow: / rule for OAI-SearchBot can remove one possible technical blocker. It cannot make an irrelevant page relevant to a user’s question, manufacture evidence, or force a citation.
For publishers, the useful question therefore changes from “Did we allow the bot?” to “Do we have a page that directly and credibly answers the kinds of questions our audience asks?”
3. The generated answer
ChatGPT search can combine current web information with generated explanation and links to sources. From a publisher’s perspective, this means visibility is not adequately described by a single traditional rank position. A brand may be mentioned without being cited; a page may be cited while the brand is not recommended; and different prompts can produce different source sets.
That is why our AI visibility tracking guide separates mentions, recommendations and citations instead of compressing them into one mysterious score.
4. Citation and referral
OpenAI says publishers that allow OAI-SearchBot can measure referral traffic from ChatGPT in analytics, and that ChatGPT automatically includes utm_source=chatgpt.com in referral URLs from search results.
Referral traffic is valuable because it is an externally observable outcome. It should still be interpreted carefully: no referral traffic does not prove that a brand was never mentioned, and a citation does not necessarily generate a click.
For a fuller measurement system, track at least three layers together: AI answer visibility, cited-source evidence, and downstream referral/conversion data.
What should robots.txt contain?
A minimal policy for a publisher that wants OAI-SearchBot to access public pages can look like this:
User-agent: OAI-SearchBot
Allow: /
Do not copy that blindly into every site. A real robots.txt file may contain path-specific restrictions, staging areas, account pages, search-result pages or other sections that should remain excluded. The objective is to give the search crawler access to the public content you actually want discoverable.
Also remember that robots.txt is a crawl-control mechanism, not a confidentiality mechanism. Sensitive content should not depend on robots rules for protection.
OAI-SearchBot and GPTBot are not interchangeable
This is one of the easiest configuration mistakes to make.
User agent
Publisher decision it relates to
OAI-SearchBot
Access for ChatGPT search discovery and search-result use
GPTBot
Potential use of web content for model training
OpenAI’s publisher FAQ specifically tells publishers to disallow GPTBot on pages they want excluded from potential training, while separately telling publishers not to block OAI-SearchBot if they want content available for ChatGPT search summaries and snippets.
So a company can make separate policy choices. That is considerably more useful than the internet’s traditional approach of treating every crawler as either angel or demon depending on which LinkedIn post was read that morning.
Why a crawler can still fail after you allow it
If OAI-SearchBot is allowed in robots.txt but the site is not being crawled as expected, inspect the infrastructure in layers:
Confirm the public URL returns a successful response without login.
Check robots.txt for user-agent-specific and wildcard rules.
Check CDN and WAF logs for blocked crawler requests.
Review bot-management policies, CAPTCHA and JavaScript challenges.
Check rate-limiting rules and 429 responses.
If your infrastructure requires IP allowlisting, use OpenAI’s published searchbot IP list rather than a manually observed address.
Verify canonical and indexability signals on the destination page.
OpenAI notes that crawler infrastructure can change, so teams should not rely only on short-term IP observations from logs. Its current guidance recommends combining user-agent identification, verified-bot programs where available, firewall allowlists, robots behavior and provider-level bot verification.
What if you do not want a page surfaced?
OpenAI’s publisher FAQ says that if it obtains the URL of a disallowed page through a third-party search provider or by crawling other pages, it may still surface only the link and page title when it has signals that the page is relevant. Publishers that do not want this should use a noindex meta tag.
There is a technical catch: OpenAI notes that its crawler must be allowed to crawl the page in order to read that meta tag.
This is another reason to think about crawl control and index/surfacing control separately rather than assuming robots.txt solves every problem.
Does allowing OAI-SearchBot improve your ChatGPT ranking?
There is no published basis for promising that.
Allowing the crawler addresses an eligibility and access condition. OpenAI says placement is not guaranteed and that ranking uses multiple factors designed around relevance and reliability. There is no documented “OAI-SearchBot bonus” for simply adding an allow rule.
A better sequence is:
Make the page accessible → make the answer useful → make claims verifiable → make the brand/entity clear → measure actual prompts and citations.
The first step is technical. The rest is editorial, entity-level and measurement work.
If you are deciding how much of this belongs to SEO versus GEO, our GEO vs SEO guide maps the overlap without pretending the two disciplines are completely separate.
How to test your own site
Start with a small evidence checklist rather than repeatedly asking ChatGPT “can you see my website?”
Fetch your live robots.txt and inspect the OAI-SearchBot policy.
Confirm important content pages are public and return successful HTTP responses.
Check that canonical and noindex directives match your intent.
Inspect CDN/WAF logs when crawler access is in doubt.
Use OpenAI’s current published searchbot IP ranges if infrastructure requires an IP allowlist.
Run a stable set of buyer-relevant prompts and record mentions, recommendations and cited URLs over time.
The last item is where crawler configuration turns into actual visibility measurement. Run 2 free MonitorMyGEO audits if you want a baseline across the AI platforms your buyers use instead of manually collecting answers in a spreadsheet.
What MonitorMyGEO measures after crawlability
Crawler access answers only “can this content potentially be reached?” It does not tell you whether your brand appears when a prospect asks an AI system for options.
MonitorMyGEO is designed for that second problem: tracking prompt-level brand visibility, competitors and cited evidence across ChatGPT, Claude, Gemini and Perplexity. You can see how the reporting works in the demo before deciding whether the measurement layer is useful for your team.
The distinction is worth keeping: AI readiness helps diagnose whether your web presence is technically prepared; AI visibility measurement observes what the systems actually return. Neither should be presented as a guarantee of inclusion or citation.
The practical takeaway
For ChatGPT search discovery, begin with what OpenAI actually documents. Keep the public pages you want discoverable accessible to OAI-SearchBot, make sure infrastructure does not silently block the crawler, use the published IP ranges where necessary, and treat crawler access as eligibility rather than a ranking lever.
Then measure outcomes. A technically crawlable site that never appears for commercially relevant questions still has a visibility problem. A cited site that receives no useful referral or conversion activity has a different problem. GEO becomes useful when those layers are measured separately instead of being collapsed into one promise about “ranking in ChatGPT.”