Google-Extended Does Not Do What You Think
A robots.txt line that reads User-agent: Google-Extended followed by Disallow: / appears on a great many sites, added in the belief that it removes them from AI Overviews. It does not, and it never did. The confusion is understandable, because the token’s name suggests a scope much broader than the one Google gives it.
What Google-Extended actually controls
Google-Extended is a robots.txt token with no crawler of its own. No request is ever made with Google-Extended in the user agent string; it is a label that other systems consult when deciding whether content already fetched by Googlebot may be used for a particular purpose.
That purpose is stated in Google’s AI features documentation, which directs readers to “limit AI training and grounding in some of Google’s other systems” by reading about Google-Extended. The operative words are other systems. Google’s position is that AI Overviews and AI Mode are part of Search, not separate products, so content crawled for Search is available to them. Google-Extended sits outside that boundary.
Why it does not stop AI Overviews
The same documentation points at a different set of controls for what appears in Search itself: nosnippet, data-nosnippet, max-snippet and noindex. These operate on preview text rather than on eligibility, and each has a cost. noindex removes the page from Search entirely. nosnippet suppresses the description in ordinary results as well as the generated one. max-snippet caps length everywhere it applies.
There is no snippet-level control that removes a page from an AI Overview while leaving its ordinary result intact. Anyone who added Google-Extended expecting that outcome has been running with a setting that does nothing for the thing they cared about, and — because Googlebot’s access was never affected — with no signal in their logs to tell them so. The general question of which bot to block, and what blocking each one costs, is the subject of should you block AI crawlers.
The fetchers most people have never heard of
Google’s user-triggered fetchers documentation, last updated 19 August 2026, lists nine fetchers that are neither Googlebot nor the AI crawlers most robots.txt files account for. Two are worth knowing by name.
Google-Agent— documented as being “used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request”. That is an agent visiting your site to do something, not to index it.Google-GeminiNotebook— which “requests individual URLs that Gemini Notebook users have provided as sources for their projects”, meaning a person pasted your URL into a research tool.
The rest of the list is older and less contentious: Feedfetcher, Google Read Aloud, Google Messages, Google Pinpoint, Google Site Verifier, Google Publisher Center and the Chrome Web Store fetcher. If you have been classifying log traffic and treating anything non-Googlebot as suspect, several of these have been misfiled, and the exercise in what’s actually crawling your site is the place to reconcile them.
“Generally ignore robots.txt” is the whole problem
The user-triggered fetchers documentation states the rule without hedging: “because the fetch was requested by a user, these fetchers generally ignore robots.txt rules.”
The reasoning is consistent with how the web has always treated user agency. A person following a link is not a crawler, and a tool fetching a URL because a person asked for it is treated the same way. The consequence is that robots.txt is not the control surface for this traffic at all — which matters because the distinction between crawl control and index control, covered in when to use robots.txt versus meta robots, collapses here: neither one applies. If a page must not be fetched by an agent acting for a user, the control is authentication, not a directive.
The control that works, and the doc that has not caught up
The mechanism that genuinely governs appearance in AI Overviews, AI Mode and Discover’s AI features is the Search generative AI setting in Search Console, which reached all sites worldwide on 31 August 2026. It is a per-property switch with Include as the default, and it is the only documented way to remove a site from those surfaces without removing it from Search.
Google’s own AI features page has not been revised to say so. It carries a last-updated date of 10 December 2025 and still directs readers to snippet controls and Google-Extended, which is a plausible reason the misconception persists: the page most likely to be read by someone asking this question predates the control that answers it. Check the last-updated stamp on any Google documentation page before acting on it, and prefer the Search Console setting over anything in robots.txt when the question is about AI surfaces rather than about crawling.