
Robots.txt is a text file placed at the root of a website that provides crawling instructions to web robots that support the Robots Exclusion Protocol. It tells crawlers which parts of a website they are allowed or disallowed to crawl.
Traditionally, robots.txt has been associated with search engine crawlers such as Googlebot and Bingbot. As AI Search has expanded, robots.txt has also become relevant to AI crawlers and bots used for search, retrieval, model development, and other automated purposes.
Website owners can use different rules for different user agents. This makes it possible, where supported by the crawler, to allow one bot while restricting another.
A compliant automated crawler generally checks a site's robots.txt file before crawling URLs covered by that file. It identifies the rules applicable to its user-agent and determines which paths it may access.
A basic robots.txt file can look like this:
In this example, the asterisk represents crawlers generally, the Allow: / directive permits crawling from the root path, and the Sitemap field identifies the location of the XML sitemap.
A robots.txt file is normally located at the root of the host it controls. For example:
Its scope matters. A robots.txt file applies to the host, protocol, and port where it is served. A file on one subdomain does not automatically establish crawling rules for every other subdomain.
For example, rules published at:
apply to the relevant www.example.com host rather than automatically controlling a separate host such as docs.example.com.
A User-agent identifies the crawler to which a group of robots.txt rules applies.
A general rule can use:
The asterisk is used to address crawlers that match that general group. Website owners can also create rules for specific supported crawler user-agent tokens.
For example:
This tells a compliant crawler matching ExampleBot not to crawl URLs under the specified path.
Two of the most recognizable robots.txt directives are Allow and Disallow.
| Directive | Purpose |
|---|---|
| User-agent | Identifies which crawler or crawler group the following rules apply to. |
| Allow | Specifies a path that the applicable crawler is permitted to crawl. |
| Disallow | Specifies a path that the applicable crawler should not crawl. |
| Sitemap | Provides the absolute URL of a sitemap or sitemap index. |
These examples are intentionally simple. Real implementations should be reviewed carefully because incorrect rules can unintentionally restrict important pages or resources.
Search engines use crawlers to discover and revisit web content. Robots.txt gives website owners a standardized mechanism for communicating crawl preferences to crawlers that respect the protocol.
For search engines, crawl accessibility is an important technical prerequisite for discovering and processing content. However, allowing a crawler does not guarantee that a page will be indexed, ranked, or displayed prominently.
The growth of generative AI has introduced additional automated crawlers operated by AI companies and AI-powered search services.
These crawlers can serve different purposes. Depending on the provider and crawler, a bot may be used for AI Search discovery, retrieval, model development, training-related crawling, or another automated function.
This distinction is important because the term AI bot should not be treated as if every crawler performs the same task.
GPTBot is an OpenAI web crawler associated with crawling content that may be used to help make OpenAI's generative AI foundation models more useful and safe.
Website owners can address GPTBot separately in robots.txt. A site that does not want GPTBot to crawl its content can use rules applicable to that user-agent.
A simplified example is:
Importantly, GPTBot should not be assumed to perform the same function as OpenAI's search crawler.
OAI-SearchBot is OpenAI's crawler used for search. OpenAI documents it separately from GPTBot, allowing website owners to manage search crawling and potential model-training crawling independently.
A website can therefore choose to allow OAI-SearchBot while applying a different policy to GPTBot.
This distinction is particularly relevant for publishers and brands that want their public content to remain accessible for supported AI Search discovery while applying a different preference to training-related crawling.
Robots.txt can therefore have implications for ChatGPT Search. OpenAI states that publishers should not block OAI-SearchBot if they want their site content to be eligible to be discovered, surfaced, cited, and linked within supported ChatGPT search experiences.
This makes crawler accessibility one technical consideration within the broader process of improving visibility in AI-powered search.
Robots.txt can affect whether supported automated systems are permitted to crawl content, which makes it relevant to AI Visibility and technical AI Search accessibility.
However, allowing an AI crawler does not guarantee that a website will be mentioned or cited in AI Answers.
AI-generated visibility can depend on many additional factors, including relevance, content quality, source selection, authority, query context, retrieval systems, and the behavior of the individual AI platform.
AI Citations occur when supported AI-generated experiences reference a website, page, or source within an answer.
Crawl accessibility can be relevant to source discovery, but allowing a bot does not guarantee that a URL will receive citations.
Teams can use Citation Monitoring to observe which domains and pages are actually being cited across monitored AI-generated answers.
With AI citation monitoring, brands can analyze their own citations, competitor citations, cited URLs, source domains, and changes over time.
Robots.txt itself is a crawler-access mechanism. Whether a particular robots.txt user-agent controls search crawling, training-related crawling, or another use depends on how the crawler operator defines and implements that user-agent.
This is why individual crawler documentation matters.
For example, OpenAI distinguishes GPTBot from OAI-SearchBot. Blocking GPTBot communicates a different preference from blocking OAI-SearchBot because the two crawlers have different documented purposes.
Robots.txt and noindex solve different problems.
Robots.txt primarily controls crawling access for compliant crawlers. A noindex directive is intended to tell supported search systems not to index a page.
Blocking a URL in robots.txt should therefore not automatically be treated as equivalent to requesting that the URL disappear from search indexes.
A crawler may need to access a page in order to discover page-level directives such as a meta robots noindex instruction.
Robots.txt and XML sitemaps also have different roles.
Robots.txt communicates crawling rules, while a sitemap helps supported crawlers discover URLs that a website wants them to know about.
A sitemap location can also be declared directly within robots.txt:
Robots.txt is not an AI visibility ranking factor or a guarantee of inclusion in AI-generated answers.
Its role is more fundamental: it can help ensure that crawlers you want to access public content are not unintentionally blocked.
If an important search crawler cannot access content because of robots.txt, CDN rules, firewall configuration, authentication, or bot protection, optimization elsewhere on the page may not solve the underlying access problem.
Technical accessibility should therefore be considered one layer of a broader LLM SEO, AEO, SEO, and Generative Engine Optimization (GEO) strategy.
A correct robots.txt configuration does not necessarily mean a crawler can successfully retrieve a page.
Access can also be affected by:
For AI Search accessibility, teams should therefore evaluate the complete path between the crawler and the content rather than checking robots.txt alone.
Robots.txt is simple in appearance but can have broad effects when rules are applied incorrectly.
Common mistakes can include:
A practical AI Search robots.txt audit can include the following steps:
Robots.txt is only one technical signal in a much larger AI Search ecosystem. Allowing a crawler tells it that it may access specified content; it does not explain whether the brand is actually visible, mentioned, cited, or preferred in relevant AI-generated answers.
Using Ansvisor's AI Search Intelligence Platform, teams can analyze what happens beyond technical accessibility by monitoring prompts, answers, visibility, citations, competitors, sources, and changes across AI-powered discovery environments.
In this broader workflow, robots.txt helps define crawler access while AI Search intelligence helps determine whether that accessible content is actually contributing to measurable visibility and business opportunities.
Robots.txt is a text file placed at the root of a website that communicates crawling rules to web crawlers that support the Robots Exclusion Protocol. It can specify which paths particular user agents are allowed or disallowed to crawl. Google for Developers
It can affect crawler accessibility, which can matter for supported AI Search systems, but allowing a crawler does not guarantee visibility or citations. For example, OpenAI states that publishers should allow OAI-SearchBot if they want content to be eligible to be discovered, surfaced, cited, and linked in supported ChatGPT search experiences.
OpenAI documents them as separate crawlers with different purposes. OAI-SearchBot is used for search, while GPTBot crawls content that may be used in training OpenAI's generative AI foundation models. Their robots.txt settings can be managed independently. OpenAI Developers
No. Robots.txt primarily controls crawling, while noindex is intended to prevent supported systems from indexing a page. A blocked crawler may be unable to access the page to read page-level indexing directives.
No. Crawler access only addresses accessibility. It does not guarantee that a page will be selected as a source, cited in an AI answer, mentioned by an AI platform, or receive traffic. AI citations and visibility should be measured separately from crawler access.
Track how your brand appears across AI platforms, understand what drives visibility, and turn insights into measurable actions.
Platform Features
Explore all features →Understand how AI platforms talk about your brand.
Discover and monitor the prompts shaping your AI visibility.
Track which sources AI platforms cite and where your brand appears.
Measure visits coming from ChatGPT, Gemini, Claude, and more.
Compare AI visibility and uncover competitive gaps and opportunities.
Turn AI Search signals into prioritized actions and executable tasks.
AI Visibility Trackers
Explore AI Visibility Platform →Track brand mentions, citations, prompts, and visibility across ChatGPT.
Monitor where and how your brand appears in Google AI Overviews.
Track your brand's visibility across Google AI Mode experiences.
Understand how your brand appears across Google Gemini responses.
Monitor your brand's presence across Microsoft Copilot answers.
Track brand mentions, citations, and visibility across Perplexity.
From AI Visibility insights to action.
Explore the complete Ansvisor platform for AI Search intelligence, optimization, and growth.
Understand, measure, and optimize your AI visibility via Ansvisor.
✓ Add brand, domains and competitors
✓ Discover prompts and growth opportunities
✓ Track your AI visibility across major AI platforms
✓ Monitor citations, mentions, and competitors
✓ Measure AI traffic and customer discovery
✓ Receive AI recommendations based on AI insights
✓ Optimize authority, trust, and content quality
✓ Create content, automate analysis & action with AI agents
Continue exploring key AI visibility concepts.
Measure and improve how often your brand appears in AI-generated answers.
Learn more →Strategies for increasing visibility in answer engines and AI summaries.
Learn more →Optimizing content for AI-powered discovery experiences.
Learn more →Understand how OpenAI retrieves and synthesizes information.
Learn more →AI-generated summaries that appear directly in Google Search.
Learn more →Explore how Perplexity cites and presents sources.
Learn more →References and sources used by AI systems to support answers.
Learn more →Measure the quality and influence of cited sources.
Learn more →How easily AI systems can discover and reuse your content.
Learn more →New terms are added regularly.
Help us improve the page or suggest a new term →
Co-founder at Ansvisor
Cihan Geyik is the co-founder of Ansvisor, an open-source AI Visibility platform for AI Search. With more than 15 years of experience in digital marketing and growth, he writes about AI visibility, AI search, AEO, GEO, citations, and answer engines. He focuses on helping brands understand and improve their presence across ChatGPT, Gemini, Perplexity, Google AI Overviews, and other AI-powered discovery platforms.
© 2026 Ansvisor. All rights reserved. Ansvisor is an open-source AI Search Intelligence Platform for AI Visibility.