AI & Infrastructure
Robots.txt file controlling access for search engine crawlers, AI crawlers, GPTBot, OAI-SearchBot, and other web bots

Robots.txt

Robots.txt is a text file that tells compliant web crawlers which parts of a website they are allowed or disallowed to crawl, including search engine crawlers and supported AI crawlers.
October 5, 2026
Cihan Geyik
Table of Content

Robots.txt is a text file placed at the root of a website that provides crawling instructions to web robots that support the Robots Exclusion Protocol. It tells crawlers which parts of a website they are allowed or disallowed to crawl.

Traditionally, robots.txt has been associated with search engine crawlers such as Googlebot and Bingbot. As AI Search has expanded, robots.txt has also become relevant to AI crawlers and bots used for search, retrieval, model development, and other automated purposes.

Website owners can use different rules for different user agents. This makes it possible, where supported by the crawler, to allow one bot while restricting another.

Crawler Requests URL → Reads robots.txt → Evaluates User-Agent Rules → Crawls or Avoids the URL

How Does Robots.txt Work?

A compliant automated crawler generally checks a site's robots.txt file before crawling URLs covered by that file. It identifies the rules applicable to its user-agent and determines which paths it may access.

A basic robots.txt file can look like this:

User-agent: * Allow: / Sitemap: https://example.com/sitemap.xml

In this example, the asterisk represents crawlers generally, the Allow: / directive permits crawling from the root path, and the Sitemap field identifies the location of the XML sitemap.

Where Is the Robots.txt File Located?

A robots.txt file is normally located at the root of the host it controls. For example:

https://example.com/robots.txt

Its scope matters. A robots.txt file applies to the host, protocol, and port where it is served. A file on one subdomain does not automatically establish crawling rules for every other subdomain.

For example, rules published at:

https://www.example.com/robots.txt

apply to the relevant www.example.com host rather than automatically controlling a separate host such as docs.example.com.

What Is a User-Agent in Robots.txt?

A User-agent identifies the crawler to which a group of robots.txt rules applies.

A general rule can use:

User-agent: *

The asterisk is used to address crawlers that match that general group. Website owners can also create rules for specific supported crawler user-agent tokens.

For example:

User-agent: ExampleBot Disallow: /private/

This tells a compliant crawler matching ExampleBot not to crawl URLs under the specified path.

Allow and Disallow in Robots.txt

Two of the most recognizable robots.txt directives are Allow and Disallow.

Directive Purpose
User-agent Identifies which crawler or crawler group the following rules apply to.
Allow Specifies a path that the applicable crawler is permitted to crawl.
Disallow Specifies a path that the applicable crawler should not crawl.
Sitemap Provides the absolute URL of a sitemap or sitemap index.

Allow All Crawling

User-agent: * Allow: /

Block a Directory

User-agent: * Disallow: /private/

Block Crawling Across the Site

User-agent: * Disallow: /

These examples are intentionally simple. Real implementations should be reviewed carefully because incorrect rules can unintentionally restrict important pages or resources.

Robots.txt and Search Engine Crawlers

Search engines use crawlers to discover and revisit web content. Robots.txt gives website owners a standardized mechanism for communicating crawl preferences to crawlers that respect the protocol.

For search engines, crawl accessibility is an important technical prerequisite for discovering and processing content. However, allowing a crawler does not guarantee that a page will be indexed, ranked, or displayed prominently.

Crawl Permission ≠ Indexing Guarantee ≠ Ranking Guarantee

Robots.txt and AI Crawlers

The growth of generative AI has introduced additional automated crawlers operated by AI companies and AI-powered search services.

These crawlers can serve different purposes. Depending on the provider and crawler, a bot may be used for AI Search discovery, retrieval, model development, training-related crawling, or another automated function.

This distinction is important because the term AI bot should not be treated as if every crawler performs the same task.

AI Crawler → Identify Purpose → Review User-Agent → Configure Access → Monitor Crawl Behavior

Robots.txt and GPTBot

GPTBot is an OpenAI web crawler associated with crawling content that may be used to help make OpenAI's generative AI foundation models more useful and safe.

Website owners can address GPTBot separately in robots.txt. A site that does not want GPTBot to crawl its content can use rules applicable to that user-agent.

A simplified example is:

User-agent: GPTBot Disallow: /

Importantly, GPTBot should not be assumed to perform the same function as OpenAI's search crawler.

Robots.txt and OAI-SearchBot

OAI-SearchBot is OpenAI's crawler used for search. OpenAI documents it separately from GPTBot, allowing website owners to manage search crawling and potential model-training crawling independently.

A website can therefore choose to allow OAI-SearchBot while applying a different policy to GPTBot.

User-agent: OAI-SearchBot Allow: / User-agent: GPTBot Disallow: /

This distinction is particularly relevant for publishers and brands that want their public content to remain accessible for supported AI Search discovery while applying a different preference to training-related crawling.

Robots.txt and ChatGPT Search

Robots.txt can therefore have implications for ChatGPT Search. OpenAI states that publishers should not block OAI-SearchBot if they want their site content to be eligible to be discovered, surfaced, cited, and linked within supported ChatGPT search experiences.

This makes crawler accessibility one technical consideration within the broader process of improving visibility in AI-powered search.

OAI-SearchBot → Search-related crawling
GPTBot → Potential foundation-model training-related crawling

Robots.txt and AI Search Visibility

Robots.txt can affect whether supported automated systems are permitted to crawl content, which makes it relevant to AI Visibility and technical AI Search accessibility.

However, allowing an AI crawler does not guarantee that a website will be mentioned or cited in AI Answers.

AI-generated visibility can depend on many additional factors, including relevance, content quality, source selection, authority, query context, retrieval systems, and the behavior of the individual AI platform.

Crawl Accessibility → Content Discovery / Retrieval → Potential Source Selection → AI Answer → Visibility & Citations

Robots.txt and AI Citations

AI Citations occur when supported AI-generated experiences reference a website, page, or source within an answer.

Crawl accessibility can be relevant to source discovery, but allowing a bot does not guarantee that a URL will receive citations.

Teams can use Citation Monitoring to observe which domains and pages are actually being cited across monitored AI-generated answers.

With AI citation monitoring, brands can analyze their own citations, competitor citations, cited URLs, source domains, and changes over time.

Does Robots.txt Control AI Training?

Robots.txt itself is a crawler-access mechanism. Whether a particular robots.txt user-agent controls search crawling, training-related crawling, or another use depends on how the crawler operator defines and implements that user-agent.

This is why individual crawler documentation matters.

For example, OpenAI distinguishes GPTBot from OAI-SearchBot. Blocking GPTBot communicates a different preference from blocking OAI-SearchBot because the two crawlers have different documented purposes.

Do not treat every AI crawler as interchangeable.

Review the crawler's owner, documented purpose, user-agent token, and official access-control guidance before creating robots.txt rules.

Robots.txt vs. Noindex

Robots.txt and noindex solve different problems.

Robots.txt primarily controls crawling access for compliant crawlers. A noindex directive is intended to tell supported search systems not to index a page.

robots.txt → Controls crawling
noindex → Controls indexing for systems that support the directive

Blocking a URL in robots.txt should therefore not automatically be treated as equivalent to requesting that the URL disappear from search indexes.

A crawler may need to access a page in order to discover page-level directives such as a meta robots noindex instruction.

Robots.txt vs. Sitemap.xml

Robots.txt and XML sitemaps also have different roles.

Robots.txt communicates crawling rules, while a sitemap helps supported crawlers discover URLs that a website wants them to know about.

robots.txt → Where crawlers may or may not crawl
sitemap.xml → URLs a site makes available for discovery

A sitemap location can also be declared directly within robots.txt:

Sitemap: https://example.com/sitemap.xml

Can Robots.txt Improve AI Visibility?

Robots.txt is not an AI visibility ranking factor or a guarantee of inclusion in AI-generated answers.

Its role is more fundamental: it can help ensure that crawlers you want to access public content are not unintentionally blocked.

If an important search crawler cannot access content because of robots.txt, CDN rules, firewall configuration, authentication, or bot protection, optimization elsewhere on the page may not solve the underlying access problem.

Technical accessibility should therefore be considered one layer of a broader LLM SEO, AEO, SEO, and Generative Engine Optimization (GEO) strategy.

Robots.txt Is Not the Only Crawler Access Layer

A correct robots.txt configuration does not necessarily mean a crawler can successfully retrieve a page.

Access can also be affected by:

  • CDN configuration;
  • web application firewalls;
  • bot protection systems;
  • CAPTCHA or JavaScript challenges;
  • authentication requirements;
  • IP-based restrictions;
  • geographic restrictions;
  • rate limiting;
  • HTTP errors; and
  • server availability.

For AI Search accessibility, teams should therefore evaluate the complete path between the crawler and the content rather than checking robots.txt alone.

robots.txt → CDN / WAF → Server Access → Page Retrieval → Content Processing

Common Robots.txt Mistakes

Robots.txt is simple in appearance but can have broad effects when rules are applied incorrectly.

Common mistakes can include:

  • accidentally blocking the entire website;
  • blocking important public directories;
  • assuming robots.txt is the same as noindex;
  • using the wrong user-agent token;
  • assuming every AI crawler has the same purpose;
  • forgetting that subdomains can require their own robots.txt configuration;
  • blocking resources required for proper page processing;
  • changing crawler rules without monitoring the result; and
  • allowing a crawler in robots.txt while blocking it at the CDN or firewall layer.

How to Audit Robots.txt for AI Search

A practical AI Search robots.txt audit can include the following steps:

  1. Locate the robots.txt file. Confirm that it is accessible from the root of the relevant host.
  2. Review global rules. Check the directives applied to the wildcard user-agent.
  3. Identify search crawlers. Determine which traditional search crawlers should access the site.
  4. Identify AI crawlers. Review the official documentation for relevant AI bots and their purposes.
  5. Separate crawler purposes. Distinguish search, retrieval, training-related, and other crawler functions where providers document them separately.
  6. Review important paths. Confirm that public pages intended for discovery are not unintentionally blocked.
  7. Check infrastructure. Review CDN, WAF, bot protection, authentication, and server responses.
  8. Verify the sitemap. Ensure the correct sitemap is discoverable where appropriate.
  9. Monitor results. Observe crawler activity, visibility, citations, and technical changes over time.

Robots.txt and AI Search Intelligence

Robots.txt is only one technical signal in a much larger AI Search ecosystem. Allowing a crawler tells it that it may access specified content; it does not explain whether the brand is actually visible, mentioned, cited, or preferred in relevant AI-generated answers.

Using Ansvisor's AI Search Intelligence Platform, teams can analyze what happens beyond technical accessibility by monitoring prompts, answers, visibility, citations, competitors, sources, and changes across AI-powered discovery environments.

Crawler Accessibility → AI Search Data → Visibility → Citations → Opportunities → Actions → Measurement

In this broader workflow, robots.txt helps define crawler access while AI Search intelligence helps determine whether that accessible content is actually contributing to measurable visibility and business opportunities.

Robots.txt, Robots Exclusion Protocol, REP, Robots File, Robots.txt File, Crawler Access File, Bot Access Rules, Crawler Directives, AI Crawler Rules, AI Bot Access Control

FAQ

Frequently asked questions.

What is robots.txt?

Robots.txt is a text file placed at the root of a website that communicates crawling rules to web crawlers that support the Robots Exclusion Protocol. It can specify which paths particular user agents are allowed or disallowed to crawl. Google for Developers

Does robots.txt affect AI Search visibility?

It can affect crawler accessibility, which can matter for supported AI Search systems, but allowing a crawler does not guarantee visibility or citations. For example, OpenAI states that publishers should allow OAI-SearchBot if they want content to be eligible to be discovered, surfaced, cited, and linked in supported ChatGPT search experiences.

What is the difference between GPTBot and OAI-SearchBot?

OpenAI documents them as separate crawlers with different purposes. OAI-SearchBot is used for search, while GPTBot crawls content that may be used in training OpenAI's generative AI foundation models. Their robots.txt settings can be managed independently. OpenAI Developers

Is blocking a page in robots.txt the same as using noindex?

No. Robots.txt primarily controls crawling, while noindex is intended to prevent supported systems from indexing a page. A blocked crawler may be unable to access the page to read page-level indexing directives.

Does allowing AI crawlers guarantee AI citations?

No. Crawler access only addresses accessibility. It does not guarantee that a page will be selected as a source, cited in an AI answer, mentioned by an AI platform, or receive traffic. AI citations and visibility should be measured separately from crawler access.

Explore Ansvisor

Everything You Need to Improve Your AI Visibility

Track how your brand appears across AI platforms, understand what drives visibility, and turn insights into measurable actions.

From AI Visibility insights to action.

Explore the complete Ansvisor platform for AI Search intelligence, optimization, and growth.

Explore the Platform
Ansvisor is an open-source and cloud-ready AI Visibility Platform that helps brands measure, understand, and optimize their brand's AI visibility across ChatGPT, Claude, Gemini, Google AI Overviews, and other AI search platforms.

Win customers from all major AI platforms

Understand, measure, and optimize your AI visibility via Ansvisor.

✓ Add brand, domains and competitors
✓ Discover prompts and growth opportunities
✓ Track your AI visibility across major AI platforms
✓ Monitor citations, mentions, and competitors
✓ Measure AI traffic and customer discovery
✓ Receive AI recommendations based on AI insights
✓ Optimize authority, trust, and content quality
✓ Create content, automate analysis & action with AI agents

Help us grow the AI Visibility Grossary

New terms are added regularly.

Help us improve the page or suggest a new term →
About the Author
Cihan Geyik

Cihan Geyik

Co-founder at Ansvisor

Cihan Geyik is the co-founder of Ansvisor, an open-source AI Visibility platform for AI Search. With more than 15 years of experience in digital marketing and growth, he writes about AI visibility, AI search, AEO, GEO, citations, and answer engines. He focuses on helping brands understand and improve their presence across ChatGPT, Gemini, Perplexity, Google AI Overviews, and other AI-powered discovery platforms.

Summarize with ChatGPT
Summarize with Claude
Summarize with Google
Summarize with Perplexity
Summarize with Grok