MonitorMyGEO uses essential storage for account and security functions. Google Analytics and Microsoft Clarity load only if you allow analytics. Read our Cookie Policy.
robots.txt for AI Search: A Practical GEO Configuration Guide
A practical 2026 guide to configuring robots.txt for Googlebot, Bingbot, OAI-SearchBot, GPTBot and PerplexityBot without confusing crawl controls with indexing or AI visibility.
Short answer:robots.txt is a crawl-control file, not an AI-ranking switch. In 2026, brands should configure it deliberately for each crawler they care about, because search discovery, AI-search discovery and model-development access can use different user agents. OpenAI says sites that want content eligible for ChatGPT search summaries and snippets should not block OAI-SearchBot. Bing and Google likewise use robots.txt as part of crawler access control. Allowing access can remove a technical barrier; it does not guarantee indexing, retrieval, citation, recommendation or ranking.
A robots.txt file lives at the root of a host, normally /robots.txt. It tells compliant crawlers which URL paths they may or may not request.
Three concepts are routinely collapsed into one: crawling is whether a crawler may request a URL; indexing or retrieval eligibility is whether information from that URL can enter or remain available to a search or retrieval system; and visibility is whether the URL or brand is actually surfaced, cited or recommended for a particular query.
robots.txt primarily operates at the first layer. It should not be described as a guaranteed way to control the other two.
Bing's current robots.txt guidance says a page prevented from crawling will not be indexed by Bing. OpenAI's publisher guidance is more nuanced: if OpenAI learns about a disallowed URL through another source, ChatGPT Atlas may still surface only its link and title when relevant. OpenAI says publishers who do not want that should use a noindex meta tag, while also noting that its crawler must be able to access the page to read that tag.
That is why a crawler policy should be designed from the outcome backward rather than copied from a random SEO template.
A simple crawler-policy model
Search discovery
Traditional search crawlers such as Googlebot and Bingbot need access to pages you want their search systems to crawl. Blocking important public pages can therefore undermine ordinary search discovery.
AI-search discovery
OpenAI documents OAI-SearchBot for search-related discovery. Its publisher FAQ says sites that want their content included in ChatGPT search summaries and snippets should make sure OAI-SearchBot is not blocked.
Perplexity separately documents PerplexityBot for its search system and states that it respects robots.txt.
Frequently asked questions
Does robots.txt control whether my brand appears in ChatGPT?
It controls crawler access for compliant user agents. OpenAI says not blocking OAI-SearchBot helps make public content eligible for ChatGPT search summaries and snippets, but placement is not guaranteed.
Should I allow OAI-SearchBot but block GPTBot?
That is a policy choice. OpenAI exposes them separately, so publishers can make different decisions about search discovery and potential model-development access.
Does Disallow: / remove a URL from every search or AI system?
No universal claim is safe. Different systems have different indexing and URL-discovery behavior; use each platform's documented removal or noindex controls when exclusion is the objective.
Should a sitemap be listed in robots.txt?
It is optional, but Bing documents the Sitemap directive as a way to point crawlers to a sitemap containing important URLs.
Does allowing AI crawlers improve GEO rankings?
There is no documented guarantee. Crawler access can remove a technical barrier to discovery, while retrieval, citation, recommendation and placement depend on additional systems and signals.
Research and editorial team at MonitorMyGEO covering Generative Engine Optimization, AI visibility, citations, measurement methodology and AI Readiness.
A practical guide to where IndexNow fits into GEO: faster URL discovery and freshness for participating search engines, without confusing notification speed with guaranteed indexing, rankings or AI citations
A practical 2026 guide to the web crawlers used by OpenAI, Google, Perplexity and Anthropic, what each bot does, and how robots.txt choices affect discovery, search and model-training access.
Model-development access
OpenAI documents GPTBot separately from OAI-SearchBot. This separation matters because a publisher can make one policy decision about ChatGPT search discovery and another about potential model-development use.
The practical point is not that every company must choose the same policy. It is that the policy should be intentional.
Example robots.txt patterns
The examples below are starting patterns, not universal prescriptions. Review them against your legal, licensing, privacy and publishing requirements before deploying them.
OpenAI's current documentation treats these as separate controls. Blocking GPTBot should therefore not be casually described as equivalent to blocking OAI-SearchBot.
Do not use robots.txt as an access-control system for secrets. The file itself is public, and crawler directives are instructions to compliant bots, not authentication.
Seven mistakes that create avoidable discovery problems
1. Blocking every unfamiliar AI user agent
A blanket policy may be consistent with some publishers' objectives. For a brand actively trying to be discoverable in AI search, however, it can conflict with that objective. OpenAI explicitly tells publishers seeking ChatGPT search inclusion not to block OAI-SearchBot. Review purpose before blocking by name.
2. Assuming User-agent: * is always the whole policy
Specific crawler groups can change how directives are interpreted. Bing says that when Bingbot finds a specific set of instructions for itself, it ignores the generic group, so general restrictions that should also apply to Bingbot need to be repeated in the Bingbot-specific group. Specific groups deserve testing rather than visual inspection alone.
3. Using robots.txt when you really mean noindex
Crawl control and indexing control are different jobs. If you need a page excluded from a search index, use the indexing controls documented by the relevant platform. Do not assume Disallow is a universal substitute for noindex.
4. Blocking assets needed to understand the page
Modern pages often depend on JavaScript, CSS, APIs and media. A crawler that receives the HTML but cannot access important supporting resources may not see the same useful page a person sees. Audit resource rules when debugging crawl or rendering problems.
5. Forgetting subdomains and hostnames
robots.txt applies per host. www.example.com, app.example.com and another subdomain can have different files and policies. Audit the actual hosts where important content lives.
6. Deploying syntax without testing representative URLs
A tiny wildcard or path mistake can affect thousands of URLs. Bing provides a robots.txt tester in Webmaster Tools specifically to evaluate whether URLs are allowed or blocked for a selected user agent. Test important templates: homepage, product page, blog article, category page and any path you intentionally restrict.
7. Treating an allow rule as a GEO strategy
Crawler access is an eligibility condition, not evidence that your content deserves retrieval or citation. Once access is healthy, the harder work remains: useful information, clear entities, original evidence, internal discovery, accurate claims and repeatable measurement. Our AI Visibility Audit Checklist covers those layers beyond robots.txt.
A practical robots.txt audit for AI visibility
Step 1: fetch the live file
Check the exact production hostname. Do not review a stale repository copy and assume the CDN is serving it unchanged.
Step 2: list the crawlers that matter to your policy
At minimum, many marketing teams will want to make explicit decisions about Googlebot, Bingbot, OAI-SearchBot, GPTBot and PerplexityBot. Other crawlers may matter depending on audience and policy. Our AI crawler guide explains the documented roles before you choose.
Step 3: test high-value URLs
For each relevant crawler, test URLs representing your main templates. Record whether the observed result matches the intended policy.
Step 4: check beyond robots.txt
A crawler can be allowed by robots.txt and still fail to access a site because of CDN bot protection, firewall rules, rate limits or authentication. OpenAI specifically advises publishers to check host and CDN access for its published searchbot IP ranges when troubleshooting OAI-SearchBot.
Step 5: inspect indexing controls
Review noindex, canonical URLs and other search controls separately. Do not infer them from robots.txt.
Step 6: verify sitemap discovery
A sitemap does not override a disallow rule, but it gives search engines a structured list of important URLs. Bing's robots.txt guidance documents the optional Sitemap: directive.
Step 7: establish a visibility baseline
Technical access is only the first checkpoint. After confirming crawler policy, measure whether the brand is actually mentioned, recommended and cited for commercially relevant prompts.
You can do that manually with a controlled prompt sheet, or use MonitorMyGEO's free audits to establish a repeatable baseline across supported AI platforms.
What should a brand allow?
There is no single robots.txt configuration that is correct for every organization.
A publisher concerned about model-development use may choose different access from a SaaS company whose primary objective is broad discoverability. A regulated company may have public marketing pages alongside account, application or customer areas that require different handling.
A defensible policy starts with four questions:
Is this content intentionally public?
Do we want this crawler's documented product to discover it?
Does that use align with our content and licensing policy?
Have we tested the resulting rules on representative URLs?
That is more durable than maintaining a giant copied list of bots whose purposes nobody on the team has reviewed.
robots.txt is necessary plumbing, not a ranking tactic
For GEO, robots.txt belongs in the technical eligibility layer. If a relevant crawler cannot access an important page, that can be a concrete blocker. Once access is available, however, there is no documented rule saying an Allow directive raises a brand's AI rank. OpenAI explicitly says search placement is not guaranteed.
The useful progression is:
crawler access → index/retrieval eligibility → useful evidence → observed mentions/citations → measurement over time
Each arrow introduces additional conditions. Skipping those conditions is how technical advice mutates into fake certainty.