AI & Infrastructure
Googlebot and Google-Extended showing Google Search crawling and Gemini training and grounding controls through robots.txt

Googlebot / Google-Extended

Googlebot is Google’s web crawler for Google Search and other Search features, while Google-Extended is a robots.txt product token that lets publishers manage certain uses of Google-crawled content for Gemini model training and grounding without affecting Google Search inclusion or ranking.
October 5, 2026
Cihan Geyik
Table of Content

Googlebot and Google-Extended are two different controls within Google's web crawling and AI ecosystem. Googlebot is Google's web crawler for Google Search, while Google-Extended is a robots.txt product token that lets publishers manage certain uses of content Google has crawled for Gemini model training and grounding.

The distinction is important because Google-Extended is not a separate web crawler. It does not have its own HTTP request user-agent string. Google states that crawling is performed using existing Google user-agent strings, while Google-Extended acts as a control token in robots.txt.

Googlebot → Google Search Crawling
Google-Extended → Gemini Training & Grounding Control

What Is Googlebot?

Googlebot is Google's primary web crawler for Google Search. It discovers and retrieves web content so Google can process it for its search systems.

Google states that crawling preferences directed at Googlebot affect Google Search, including Search features, as well as products such as Google Images, Google Video, Google News, and Discover.

For website owners, Googlebot accessibility is therefore a fundamental component of traditional search crawlability and Crawler Access.

Googlebot → Crawl → Google Search Systems → Search Features

What Is Google-Extended?

Google-Extended is a standalone robots.txt product token provided by Google.

It allows web publishers to manage whether content Google crawls from their websites may be used for two broad purposes:

  • training future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini; and
  • grounding in Gemini Apps and Grounding with Google Search in Vertex AI.

Grounding refers to providing content from the Google Search index to a model at prompt time to improve factuality and relevance.

Google-Extended is a control token, not a separate crawler. There is no independent Google-Extended HTTP crawler user-agent.

Googlebot vs. Google-Extended

Attribute Googlebot Google-Extended
Type Web crawler robots.txt product token
Separate HTTP user-agent Yes No
Google Search crawling Yes No separate crawling process
Gemini model training control Not its primary robots.txt purpose Yes
Gemini grounding control Not the dedicated control token Yes
Affects Google Search inclusion Googlebot access can affect crawling No
Google Search ranking signal Not applicable as a simple allow/block ranking signal No

Google-Extended Is Not a Separate AI Crawler

One of the most common misconceptions is describing Google-Extended as Google's equivalent of a standalone AI crawler.

Google explicitly states that Google-Extended does not have a separate HTTP request user-agent string.

Instead, Google crawls using its existing Google user agents. The Google-Extended token is then used in robots.txt as a control over specific downstream uses of the crawled content.

Existing Google Crawlers → Website Content → Google Systems
Google-Extended → Controls Certain Gemini Uses of That Content

This makes Google-Extended structurally different from crawlers such as OAI-SearchBot, GPTBot, ClaudeBot, or Claude-SearchBot.

How to Allow Googlebot

A website can explicitly allow Googlebot to crawl its public content with:

User-agent: Googlebot Allow: /

These directives are defined in the site's robots.txt file.

How to Block Googlebot

A site-wide Googlebot crawling restriction can be expressed as:

User-agent: Googlebot Disallow: /

Because Googlebot is used for Google Search crawling, blocking it can prevent Google from crawling affected content and can therefore have major consequences for Search discovery and indexing.

Do not block Googlebot simply because you want to restrict Gemini-related use. Google provides Google-Extended for that separate control.

How to Allow Google-Extended

A publisher that wants to permit the uses controlled by Google-Extended can use:

User-agent: Google-Extended Allow: /

This token applies to the Gemini-related uses Google documents for Google-Extended.

How to Block Google-Extended

A publisher can restrict the uses controlled by Google-Extended with:

User-agent: Google-Extended Disallow: /

This does not mean that a separate Google-Extended crawler stops visiting the website, because Google-Extended is not a separate crawler.

Instead, the directive controls whether content crawled by Google can be used for the Gemini training and grounding purposes covered by the token.

Can You Allow Googlebot but Block Google-Extended?

Yes. This is one of the most important capabilities of Google's crawler controls.

A publisher can allow Googlebot for Google Search while restricting the uses governed by Google-Extended:

User-agent: Googlebot Allow: / User-agent: Google-Extended Disallow: /
Googlebot Allowed → Google Search Crawling Allowed
Google-Extended Blocked → Covered Gemini Uses Restricted

Google states that Google-Extended does not affect whether a site is included in Google Search.

Does Blocking Google-Extended Hurt Google Rankings?

No. Google explicitly states that Google-Extended is not used as a ranking signal in Google Search.

Google also states that Google-Extended does not affect a site's inclusion in Google Search.

Google-Extended blocked ≠ Google Search blocked.
Google-Extended blocked ≠ Google Search ranking penalty.

This distinction allows publishers to make a separate decision about Gemini-related use without using Googlebot as a blunt control.

Google-Extended and Gemini Model Training

Google-Extended can be used to manage whether content crawled from a website may be used to train future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini.

This makes Google-Extended partly a model-training governance mechanism.

Google Crawling → Website Content → Google-Extended Policy → Future Gemini Model Training Use

The distinction between crawling and downstream model use is important: Google-Extended controls the specified use of crawled content rather than representing a separate crawler performing the crawl.

Google-Extended and Gemini Grounding

Google-Extended also controls specified grounding uses in Gemini Apps and Grounding with Google Search in Vertex AI.

In this context, grounding means providing content from the Google Search index to the model at prompt time to improve the factuality and relevance of its response.

User Prompt → Gemini → Google Search Grounding → Relevant Search Index Content → Grounded Response

This is different from model training. Training changes or contributes to model development, while grounding provides external information to the model during response generation.

Googlebot and Google AI Overviews

Google states that Googlebot crawling preferences affect Google Search and all Google Search features.

This distinction matters for Google AI Overviews, which are part of the broader Google Search experience.

Publishers should therefore avoid assuming that they need a separate "AI Overviews crawler" in order for content to participate in Google Search's AI features.

Google AI Overviews are part of Google Search. Googlebot remains relevant to the underlying Search crawling and indexing ecosystem.

Googlebot and Google AI Mode

Google AI Mode is also a Google Search experience rather than a completely independent web index operated through a crawler called "AI Mode Bot."

For publishers, this means the technical foundations of Google Search remain relevant when optimizing for visibility across Google's AI Search experiences.

Teams measuring performance can use Ansvisor's Google AI Mode Visibility Tracker to monitor actual visibility rather than inferring performance from crawler access alone.

Googlebot and AI Search Indexing

Googlebot plays a fundamental role in Google's discovery and crawling pipeline, which connects to AI Search Indexing when Google's search infrastructure supports AI-generated Search experiences.

However, crawling, indexing, retrieval, ranking, and AI answer generation remain distinct stages.

Googlebot Crawl → Google Search Index → Retrieval & Ranking → Search / AI Search Experience

A page being crawlable does not guarantee that it will be indexed, retrieved, surfaced, cited, or clicked.

Does Google-Extended Control Google Search Indexing?

No. Google explicitly states that Google-Extended does not affect a site's inclusion in Google Search.

It should therefore not be treated as an indexing directive equivalent to blocking Googlebot or applying a noindex directive.

Googlebot → Search Crawling
Indexing Controls → Search Index Eligibility
Google-Extended → Specified Gemini Training & Grounding Uses

Google-Extended vs. GPTBot

Google-Extended and GPTBot can both appear in discussions about AI model training controls, but they work differently.

Control Type Primary Role
Google-Extended robots.txt product token Controls specified Gemini training and grounding uses of Google-crawled content
GPTBot Web crawler Crawls content that may be used in training OpenAI generative AI foundation models

In other words, GPTBot identifies a crawler, while Google-Extended identifies a downstream-use control.

Google-Extended vs. ClaudeBot

The same structural distinction applies when comparing Google-Extended with ClaudeBot.

ClaudeBot is an Anthropic crawler identity associated with potential model training use. Google-Extended, by contrast, does not perform its own separate HTTP crawling.

Not every AI robots.txt token represents a crawler. Understanding whether a token controls crawling or downstream content use is essential when creating AI crawler policies.

Googlebot and AI Crawler Accessibility

Googlebot access is one component of AI Crawler Accessibility when Google's Search infrastructure is relevant to AI-generated Search experiences.

A robots.txt rule is only one layer. Actual crawling can also be affected by:

  • CDN configuration;
  • web application firewalls;
  • bot protection;
  • authentication;
  • IP restrictions;
  • rate limiting;
  • server availability;
  • HTTP errors;
  • redirect behavior; and
  • resource accessibility.
robots.txt → Infrastructure → Googlebot → Content Retrieval → Search Processing

How to Verify Googlebot

HTTP user-agent strings can be spoofed, so seeing "Googlebot" in a request is not by itself proof that the request came from Google.

Google publishes technical guidance and IP range information for verifying legitimate Google crawler traffic.

Engineering teams should use Google's current official verification methods rather than relying solely on user-agent matching.

Can You Verify Google-Extended Traffic?

There is no separate Google-Extended HTTP crawler traffic to identify, because Google-Extended does not have its own HTTP request user-agent string.

Existing Google user-agent strings perform the crawling, while the Google-Extended robots.txt token controls the specified downstream uses.

Searching server logs for a dedicated "Google-Extended crawler" misunderstands how Google-Extended works.

Does Allowing Googlebot Guarantee AI Visibility?

No. Googlebot accessibility enables Google to crawl eligible content, but it does not guarantee visibility in Google Search or its AI-generated experiences.

Crawling is only an early stage in the process.

Crawlable ≠ Indexed ≠ Retrieved ≠ Ranked ≠ Visible in AI Results

Actual AI Visibility must therefore be measured independently.

Does Allowing Google-Extended Improve Google AI Visibility?

Google does not describe Google-Extended as a Google Search ranking signal. In fact, Google explicitly states that Google-Extended is not used as a ranking signal in Google Search.

Publishers should therefore avoid treating Allow: / for Google-Extended as an SEO ranking tactic.

Google-Extended is a content-use control, not a Google Search ranking lever.

Common Googlebot and Google-Extended Mistakes

  • describing Google-Extended as a separate crawler;
  • searching server logs for a dedicated Google-Extended HTTP user-agent;
  • blocking Googlebot when the actual goal is to restrict specified Gemini uses;
  • assuming blocking Google-Extended removes a site from Google Search;
  • assuming blocking Google-Extended creates a Google Search ranking penalty;
  • assuming allowing Google-Extended creates a ranking benefit;
  • confusing Gemini model training with Gemini grounding;
  • assuming Google AI Overviews use an entirely separate crawler and web index;
  • confusing crawler accessibility with guaranteed AI visibility; and
  • treating every robots.txt AI-related token as if it represents an independent bot.

How to Audit Googlebot and Google-Extended

  1. Review Googlebot access. Confirm that important public URLs intended for Google Search are not unintentionally blocked.
  2. Review Google-Extended separately. Decide whether the organization wants to permit the Gemini training and grounding uses governed by this token.
  3. Do not treat the controls as interchangeable. Googlebot controls crawling for Search; Google-Extended governs specified downstream Gemini uses.
  4. Review infrastructure access. Check CDN, WAF, bot protection, HTTP responses, and server availability.
  5. Verify legitimate Googlebot traffic. Use Google's current crawler verification guidance rather than trusting the user-agent string alone.
  6. Review important subdomains. Make sure crawler policy matches the intended configuration across public hosts.
  7. Measure actual AI Search visibility. Do not infer AI Overview or AI Mode performance from crawler settings.

Googlebot, Google-Extended, and LLM SEO

Understanding the difference between Googlebot and Google-Extended is an important technical component of LLM SEO, AEO, and Generative Engine Optimization (GEO).

Technical optimization requires separating three questions:

Can Google crawl the content?
Can the content participate in Google Search?
How may Google use that content across Gemini experiences?

These questions involve related infrastructure but different controls.

From Googlebot Access to Google AI Search Performance

Crawler configuration establishes technical eligibility. Performance measurement shows what actually happens in AI Search.

Ansvisor's Google AI Overviews Rank Tracker helps teams monitor their visibility in Google AI Overviews.

The Google AI Mode Visibility Tracker provides visibility measurement for Google's AI Mode experience.

Ansvisor's broader AI Search Intelligence Platform connects AI visibility, prompts, citations, competitors, opportunities, and actions across multiple AI Search platforms.

Googlebot Access → Google Search Ecosystem → AI Search Visibility → Citations → Opportunities → Actions
Googlebot, Google Bot, Google Search Crawler, Google Web Crawler, Google-Extended, Google Extended, Google AI Crawler Control, Gemini Crawler Control, Gemini Robots.txt Control, Google AI Training Control

FAQ

Frequently asked questions.

What is Googlebot?

Googlebot is Google's primary web crawler for Google Search. Google states that Googlebot crawling preferences affect Google Search and its Search features, as well as products including Google Images, Google Video, Google News, and Discover.

What is Google-Extended?

Google-Extended is a standalone robots.txt product token, not a separate crawler. It lets publishers manage whether Google-crawled content may be used for training future Gemini models and for specified grounding uses in Gemini Apps and Vertex AI. Google for Developers

Is Google-Extended a web crawler like Googlebot?

No. Google explicitly states that Google-Extended has no separate HTTP request user-agent string. Crawling takes place through existing Google user agents; Google-Extended is used as a robots.txt control token.

Can I block Google-Extended without blocking Google Search?

Yes. Google states that Google-Extended does not affect whether a site is included in Google Search. A publisher can therefore allow Googlebot while separately disallowing Google-Extended.

Does blocking Google-Extended hurt Google Search rankings?

No. Google explicitly states that Google-Extended is not used as a ranking signal in Google Search and does not affect inclusion in Google Search.

Explore Ansvisor

Everything You Need to Improve Your AI Visibility

Track how your brand appears across AI platforms, understand what drives visibility, and turn insights into measurable actions.

From AI Visibility insights to action.

Explore the complete Ansvisor platform for AI Search intelligence, optimization, and growth.

Explore the Platform
Ansvisor is an open-source and cloud-ready AI Visibility Platform that helps brands measure, understand, and optimize their brand's AI visibility across ChatGPT, Claude, Gemini, Google AI Overviews, and other AI search platforms.

Win customers from all major AI platforms

Understand, measure, and optimize your AI visibility via Ansvisor.

✓ Add brand, domains and competitors
✓ Discover prompts and growth opportunities
✓ Track your AI visibility across major AI platforms
✓ Monitor citations, mentions, and competitors
✓ Measure AI traffic and customer discovery
✓ Receive AI recommendations based on AI insights
✓ Optimize authority, trust, and content quality
✓ Create content, automate analysis & action with AI agents

Help us grow the AI Visibility Grossary

New terms are added regularly.

Help us improve the page or suggest a new term →
About the Author
Cihan Geyik

Cihan Geyik

Co-founder at Ansvisor

Cihan Geyik is the co-founder of Ansvisor, an open-source AI Visibility platform for AI Search. With more than 15 years of experience in digital marketing and growth, he writes about AI visibility, AI search, AEO, GEO, citations, and answer engines. He focuses on helping brands understand and improve their presence across ChatGPT, Gemini, Perplexity, Google AI Overviews, and other AI-powered discovery platforms.

Summarize with ChatGPT
Summarize with Claude
Summarize with Google
Summarize with Perplexity
Summarize with Grok