An AI crawler reads what your server returns to it. robots.txt does not change that: it is a public file of requests that cooperating crawlers choose to follow. Whether a page is indexed, and how its content is used, are separate questions with separate controls.
Most confusion about AI crawlers comes from treating four different questions as one. Keep them apart:
- Access: can an agent fetch the page? If your server returns the page to an anonymous request, any client can fetch it. Only authentication and server- or network-level rules actually stop a fetch.
- Permission: what does robots.txt ask? It states, crawler by crawler, which paths you would prefer not to be fetched. Cooperating crawlers follow it. It is not access control.
- Indexing: will the page appear in a search index? The search engine decides, and you influence it with
noindex— not with robots.txt. - Use: what happens to the content? Search results, AI answers, model training, an open dataset. Each operator documents its own controls, usually a separate robots.txt token for each purpose.
Each control in this article answers one of these questions. Most mistakes come from expecting it to answer the other three.
A crawler is an HTTP client. It requests a URL and receives exactly what your server sends back to that request: the HTML, and any images, scripts or files it then asks for. If a page loads without signing in, every client can fetch it — including clients that never read robots.txt.
Not every agent runs a page's JavaScript, and operators rarely document which of their agents do. The dependable assumption is that the text you want read must be in the HTML your server sends, through server-rendered pages or static generation — not inserted later by scripts in the browser.
Content behind authentication, or on a private network, is not available to an anonymous request. The protection has to be on the server: a paywall that sends the full text in the HTML and hides it with a script has already given the text to anyone who reads the HTML.
A CDN, a firewall, bot-management rules or rate limits can refuse a request before your site responds at all. Those rules are enforced, but they are invisible to anyone reading your robots.txt — including you, months later. They can also have side effects: Anthropic notes that blocking its IP addresses "may not work correctly or persistently guarantee an opt-out", because it stops its crawler from reading your robots.txt.
robots.txt is a plain text file at the root of a host (/robots.txt). Its format is standardised as the Robots Exclusion Protocol, RFC 9309, published by the IETF in September 2022. It is made of groups: one or more User-agent lines naming crawlers, followed by Allow and Disallow rules for URL paths.
The standard is explicit about what the file is not: "These rules are not a form of access authorization."
Under RFC 9309, a crawler:
- identifies itself by a product token — the name its operator documents, such as
GPTBot; - looks for the group whose
User-agentline matches that token, ignoring case (several groups for the same token are merged); - uses the
*group only if no group matches its token — and if there is no*group either, no rules apply; - within its group, applies the most specific matching rule — the longest path. If an
Allowand aDisalloware equally specific,Allowshould win.
Step 3 is where policies usually go wrong:
User-agent: *
Disallow: /drafts/
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Disallow: /drafts/
GPTBot matches its own group and is asked not to fetch anything. OAI-SearchBot matches its own group too — so the * group no longer applies to it, which is why /drafts/ has to be repeated there. Every other crawler falls back to *. A named group replaces the * rules for that crawler; it does not add to them.
Google's interpretation follows the same model and documents its specifics: * and $ wildcards in paths, a 500 KiB size limit, and a cache of up to 24 hours.
- A shortened or guessed token. Anthropic's crawler is
ClaudeBot. A crawler looks for the group that matches its own product token, so a group forUser-agent: Claudeis not a group forClaudeBot— it falls back to*. Copy each token exactly as its operator documents it. - Assuming one token covers an operator. OpenAI, Anthropic and Perplexity each document several agents with different purposes. Disallowing
GPTBotdoes nothing toOAI-SearchBot. - Forgetting that a named group drops the
*rules — the example above. - Disallowing a page you want removed from search. A crawler that may not fetch the page never sees its
noindex(below). - One host, one file. Rules are read from the host that serves them; Anthropic asks site owners to opt out on "every subdomain".
- A broken robots.txt. RFC 9309 tells crawlers to treat a server error (5xx) on robots.txt as a complete disallow, while a "not found" (4xx) lets them access any resource. Google documents its own handling of each case.
The rules are requests. Google puts it plainly: "it's up to the crawler to obey them. While Googlebot and other respectable web crawlers obey the instructions in a robots.txt file, other crawlers might not." The file is also public: a Disallow line for a private path tells every reader where that path is.
Google states that "a page that's disallowed in robots.txt can still be indexed if linked to from other sites" — the URL, and text such as the anchor text of links to it, can still appear. To keep a page out of Google's index, it recommends noindex — and noindex only works if the page is not blocked by robots.txt, because the crawler has to fetch the page to read it.
A rule applies to fetches made after a crawler has read it. It cannot recall copies made before the rule existed, and it says nothing about copies of your content published on other sites — those are governed by those sites, not yours.
Some operators document agents that fetch a page because a user asked their assistant to. For these, robots.txt is not a guarantee: OpenAI says of ChatGPT-User that "robots.txt rules may not apply", and Perplexity says Perplexity-User "generally ignores robots.txt rules". Anthropic, by contrast, documents that all three of its agents, including Claude-User, honour robots.txt.
Each activity has its own control:
| Activity | What it means | What governs it |
|---|---|---|
| Fetching | An agent requests a URL and receives what the server returns | Authentication and server or network rules enforce it; robots.txt only asks |
| Indexing | A search engine stores the page so it can appear in results | noindex (meta tag or X-Robots-Tag header), on a page that stays crawlable |
| Search features | Snippets and AI features shown in search results | For Google Search: nosnippet, data-nosnippet, max-snippet, noindex |
| Model training | Collected content is used to train models | The operator's documented token in robots.txt, where one exists |
| AI search | An AI product surfaces or cites pages in its answers | The operator's documented search token in robots.txt |
| User-initiated fetching | An assistant fetches a page because a user asked it to | The operator's documentation — robots.txt may not apply |
| Open datasets | Pages are archived in a public dataset others can use | The dataset crawler's token in robots.txt |
The table below lists the agents that the operators document, with their purpose and their stance on robots.txt in the operators' own terms. It was checked against each operator's page on 1 October 2026; these pages change, so check the source before relying on a row.
| Token | Operator | Documented purpose | robots.txt, per the operator |
|---|---|---|---|
Googlebot | Crawls for Google Search | Obeys robots.txt | |
Google-Extended | Not a crawler: controls whether content Google crawls may be used to train Gemini models and for grounding | A control token only; "does not impact a site's inclusion in Google Search" | |
Applebot | Apple | Powers search features in Spotlight, Siri and Safari | Respects rules for Applebot; follows Googlebot rules if Applebot isn't named |
Applebot-Extended | Apple | Not a crawler: opts content out of training Apple's generative foundation models | "Does not crawl webpages"; disallowing it doesn't remove pages from Apple's search results |
GPTBot | OpenAI | Crawls content that may be used to train OpenAI's generative AI foundation models | Disallowing it indicates content should not be used in training |
OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT's search features | Opted-out sites aren't shown in ChatGPT search answers, but "can still appear as navigational links" |
ChatGPT-User | OpenAI | Visits pages for user actions in ChatGPT and Custom GPTs | "robots.txt rules may not apply" |
OAI-AdsBot | OpenAI | Validates the safety of pages submitted as ads on ChatGPT | Not stated |
ClaudeBot | Anthropic | Collects web content "that could potentially contribute to" training Anthropic's generative AI models | Honours robots.txt |
Claude-SearchBot | Anthropic | Improves search result quality for Claude users | Honours robots.txt |
Claude-User | Anthropic | Accesses sites when Claude users ask questions | Honours robots.txt |
PerplexityBot | Perplexity | Surfaces and links websites in Perplexity's search results; "not used to crawl content for AI foundation models" | Perplexity recommends allowing it in robots.txt |
Perplexity-User | Perplexity | Visits pages when users ask questions; not used for crawling or training | "Generally ignores robots.txt rules" |
CCBot | Common Crawl | Builds Common Crawl's open repository of web crawl data | Respects robots.txt |
Googlebot and Applebot crawl for search. For the AI features in Google Search, Google lists Search controls — nosnippet, data-nosnippet, max-snippet and noindex — and points to Google-Extended separately, for training and grounding in some of its other systems.
GPTBot and ClaudeBot collect content that their operators document may be used for model training. Disallowing them in robots.txt is the opt-out each operator documents, and it applies to future crawling.
OAI-SearchBot, Claude-SearchBot and PerplexityBot fetch pages so that an AI product can surface or cite them. Blocking one is a choice about appearing in that product, not about training.
ChatGPT-User, Claude-User and Perplexity-User fetch a page when a person asks the assistant to. Their operators document different positions on robots.txt, set out above.
Google-Extended and Applebot-Extended are not crawlers. Google states that Google-Extended "doesn't have a separate HTTP request user agent string": crawling is done with Google's existing user agents, and the token is used "in a control capacity". Apple states that Applebot-Extended "does not crawl webpages". Neither token affects inclusion in its operator's search results.
"Public" means served to an anonymous request. Public content can be read by any agent, including agents that ignore robots.txt, and copies of it can circulate. If something must not be read by a crawler, robots.txt is the wrong tool: put it behind authentication, or don't publish it.
| Goal | Control | What it does | What it doesn't do |
|---|---|---|---|
| Keep content private | Authentication | Refuses every request without access | Undo copies made while it was public |
| Keep a page out of Google Search | noindex (meta tag or X-Robots-Tag) | Tells Google not to index the page once it is crawled | Work if robots.txt blocks the page |
| Limit how Google Search shows a page, AI features included | nosnippet, data-nosnippet, max-snippet | Restricts the text Google may show | Remove the page from the index |
| Opt out of an operator's model training | That operator's documented token in robots.txt | Records your choice for its future crawling | Reach content already collected, or other operators |
| Stay out of an AI search product | That operator's search token in robots.txt | Asks its search crawler not to fetch | Cover user-initiated fetchers that may not apply robots.txt |
| Reduce crawler load | robots.txt rules for cooperating crawlers | Asks them to fetch less | Enforce anything; Crawl-delay isn't part of RFC 9309, and support varies by operator |
| Stop unwanted traffic | Firewall, CDN or rate-limit rules | Refuses the request | Show up in robots.txt, or leave every agent able to read your rules |
Start from the questions, not from a list of bot names: do you want to be in search indexes, in AI search products, in model training, available to assistants fetching on a user's behalf? Each answer maps to different tokens and controls. This is the part of technical and AI-search SEO that is a business decision before it is a technical one — there is no correct default for every site.
An example, not a recommendation for every site: a policy that stays open to search and AI search, and opts out of the training and dataset uses that these operators document a token for.
# Everyone else, including search and AI search crawlers: allowed
User-agent: *
Allow: /
# Training and open datasets: opted out, by each operator's documented token
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
Sitemap: https://www.example.com/sitemap.xml
Consecutive User-agent lines share one group, so all five tokens receive the same rule. Two of them, Google-Extended and Applebot-Extended, are control tokens rather than crawlers: the rule records a usage preference, not a fetch restriction. CCBot builds an open dataset that others use for their own purposes, which is why it sits in this group.
To keep a single page out of search results, leave it crawlable and mark the page itself:
<meta name="robots" content="noindex">
For files without HTML, such as PDFs, send the same instruction as an HTTP response header:
X-Robots-Tag: noindex
To keep a page in Google Search but stop its text being shown as a snippet — which Google documents as also applying to its AI features — use nosnippet instead:
<meta name="robots" content="nosnippet">
The AI Visibility Check reads one public page and your robots.txt, and reports what your rules state for each agent in its register: allowed or disallowed, which group applied, and the operator's documented role and stance on robots.txt. It also reports the page's meta robots and X-Robots-Tag.
The register is ELIXIR's record of the operators' documentation, last checked on 27 September 2026: thirteen entries — two search crawlers, nine AI crawlers and fetchers, and two usage-control tokens. It is a reading of documentation, not an observation of crawler behaviour, and it is not exhaustive: while this article was being checked, OpenAI's page listed an agent, OAI-AdsBot, that the register does not include.
The tool states its own limit, and it applies to any reading of robots.txt: "Allowed doesn't mean a system will use or cite the page; disallowed doesn't make it invisible — and blocking at a CDN or firewall can't be seen here."
For the basics around it — whether robots.txt answers at all, the sitemap it declares, and noindex on a page — check your robots.txt, sitemap and noindex settings.
Tokens, purposes and policies change, and new agents appear. Re-read the operators' pages when you review your policy. Changes also take time to apply: Google generally caches robots.txt for up to 24 hours, and OpenAI says that for OAI-SearchBot it can take about 24 hours for its systems to adjust to a robots.txt update.
Checked on 1 October 2026.
- RFC 9309 — Robots Exclusion Protocol (IETF, September 2022)
- Google Search Central — Introduction to robots.txt
- Google Search Central — How Google interprets the robots.txt specification
- Google Search Central — Block Search indexing with noindex
- Google Search Central — Robots meta tag, data-nosnippet and X-Robots-Tag
- Google Search Central — AI features and your website
- Google Search Central — Google's common crawlers (Google-Extended)
- Google Search Central — Googlebot
- OpenAI — Overview of OpenAI crawlers
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity — Perplexity crawlers
- Common Crawl — CCBot
- Apple — About Applebot
