Skip to content

Digital transformation

What AI crawlers can actually read on a website — and what robots.txt does and doesn't control

  • 13 min read
  • By ELIXIR Creative
Diagram: a visitor reaches a website and its server. Two crawlers approach from the right, and a dashed gold line marked robots.txt, “a request, not a lock”, stands in front of the server: one crawler stops at a Disallow, the other is allowed through to the server. Below, a sequence of four steps: access, permission, indexing, use.

AI crawlers read what your server publicly returns. robots.txt is a public request that cooperating crawlers follow — not access control, and not noindex. What each control actually governs, with sources.

An AI crawler reads what your server returns to it. robots.txt does not change that: it is a public file of requests that cooperating crawlers choose to follow. Whether a page is indexed, and how its content is used, are separate questions with separate controls.

The short answer

Most confusion about AI crawlers comes from treating four different questions as one. Keep them apart:

  • Access: can an agent fetch the page? If your server returns the page to an anonymous request, any client can fetch it. Only authentication and server- or network-level rules actually stop a fetch.
  • Permission: what does robots.txt ask? It states, crawler by crawler, which paths you would prefer not to be fetched. Cooperating crawlers follow it. It is not access control.
  • Indexing: will the page appear in a search index? The search engine decides, and you influence it with noindex — not with robots.txt.
  • Use: what happens to the content? Search results, AI answers, model training, an open dataset. Each operator documents its own controls, usually a separate robots.txt token for each purpose.

Each control in this article answers one of these questions. Most mistakes come from expecting it to answer the other three.

What a crawler can technically access

Anything your server returns publicly

A crawler is an HTTP client. It requests a URL and receives exactly what your server sends back to that request: the HTML, and any images, scripts or files it then asks for. If a page loads without signing in, every client can fetch it — including clients that never read robots.txt.

The initial HTML — and why server-rendered text matters

Not every agent runs a page's JavaScript, and operators rarely document which of their agents do. The dependable assumption is that the text you want read must be in the HTML your server sends, through server-rendered pages or static generation — not inserted later by scripts in the browser.

What a crawler cannot reach

Content behind authentication, or on a private network, is not available to an anonymous request. The protection has to be on the server: a paywall that sends the full text in the HTML and hides it with a script has already given the text to anyone who reads the HTML.

What can block a crawler before robots.txt is read

A CDN, a firewall, bot-management rules or rate limits can refuse a request before your site responds at all. Those rules are enforced, but they are invisible to anyone reading your robots.txt — including you, months later. They can also have side effects: Anthropic notes that blocking its IP addresses "may not work correctly or persistently guarantee an opt-out", because it stops its crawler from reading your robots.txt.

What robots.txt is

robots.txt is a plain text file at the root of a host (/robots.txt). Its format is standardised as the Robots Exclusion Protocol, RFC 9309, published by the IETF in September 2022. It is made of groups: one or more User-agent lines naming crawlers, followed by Allow and Disallow rules for URL paths.

The standard is explicit about what the file is not: "These rules are not a form of access authorization."

How a crawler finds the rules that apply to it

Under RFC 9309, a crawler:

  1. identifies itself by a product token — the name its operator documents, such as GPTBot;
  2. looks for the group whose User-agent line matches that token, ignoring case (several groups for the same token are merged);
  3. uses the * group only if no group matches its token — and if there is no * group either, no rules apply;
  4. within its group, applies the most specific matching rule — the longest path. If an Allow and a Disallow are equally specific, Allow should win.

Step 3 is where policies usually go wrong:

User-agent: *
Disallow: /drafts/

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /drafts/

GPTBot matches its own group and is asked not to fetch anything. OAI-SearchBot matches its own group too — so the * group no longer applies to it, which is why /drafts/ has to be repeated there. Every other crawler falls back to *. A named group replaces the * rules for that crawler; it does not add to them.

Google's interpretation follows the same model and documents its specifics: * and $ wildcards in paths, a 500 KiB size limit, and a cache of up to 24 hours.

Common mistakes in AI crawler rules

  • A shortened or guessed token. Anthropic's crawler is ClaudeBot. A crawler looks for the group that matches its own product token, so a group for User-agent: Claude is not a group for ClaudeBot — it falls back to *. Copy each token exactly as its operator documents it.
  • Assuming one token covers an operator. OpenAI, Anthropic and Perplexity each document several agents with different purposes. Disallowing GPTBot does nothing to OAI-SearchBot.
  • Forgetting that a named group drops the * rules — the example above.
  • Disallowing a page you want removed from search. A crawler that may not fetch the page never sees its noindex (below).
  • One host, one file. Rules are read from the host that serves them; Anthropic asks site owners to opt out on "every subdomain".
  • A broken robots.txt. RFC 9309 tells crawlers to treat a server error (5xx) on robots.txt as a complete disallow, while a "not found" (4xx) lets them access any resource. Google documents its own handling of each case.

What robots.txt does not control

It is not access control or security

The rules are requests. Google puts it plainly: "it's up to the crawler to obey them. While Googlebot and other respectable web crawlers obey the instructions in a robots.txt file, other crawlers might not." The file is also public: a Disallow line for a private path tells every reader where that path is.

It does not keep a page out of search results

Google states that "a page that's disallowed in robots.txt can still be indexed if linked to from other sites" — the URL, and text such as the anchor text of links to it, can still appear. To keep a page out of Google's index, it recommends noindex — and noindex only works if the page is not blocked by robots.txt, because the crawler has to fetch the page to read it.

It does not reach content already collected, or copied elsewhere

A rule applies to fetches made after a crawler has read it. It cannot recall copies made before the rule existed, and it says nothing about copies of your content published on other sites — those are governed by those sites, not yours.

It does not bind agents that don't follow it — or fetches its operator says it may not apply to

Some operators document agents that fetch a page because a user asked their assistant to. For these, robots.txt is not a guarantee: OpenAI says of ChatGPT-User that "robots.txt rules may not apply", and Perplexity says Perplexity-User "generally ignores robots.txt rules". Anthropic, by contrast, documents that all three of its agents, including Claude-User, honour robots.txt.

Crawling, indexing, training and answering are different things

Each activity has its own control:

ActivityWhat it meansWhat governs it
FetchingAn agent requests a URL and receives what the server returnsAuthentication and server or network rules enforce it; robots.txt only asks
IndexingA search engine stores the page so it can appear in resultsnoindex (meta tag or X-Robots-Tag header), on a page that stays crawlable
Search featuresSnippets and AI features shown in search resultsFor Google Search: nosnippet, data-nosnippet, max-snippet, noindex
Model trainingCollected content is used to train modelsThe operator's documented token in robots.txt, where one exists
AI searchAn AI product surfaces or cites pages in its answersThe operator's documented search token in robots.txt
User-initiated fetchingAn assistant fetches a page because a user asked it toThe operator's documentation — robots.txt may not apply
Open datasetsPages are archived in a public dataset others can useThe dataset crawler's token in robots.txt

The table below lists the agents that the operators document, with their purpose and their stance on robots.txt in the operators' own terms. It was checked against each operator's page on 1 October 2026; these pages change, so check the source before relying on a row.

TokenOperatorDocumented purposerobots.txt, per the operator
GooglebotGoogleCrawls for Google SearchObeys robots.txt
Google-ExtendedGoogleNot a crawler: controls whether content Google crawls may be used to train Gemini models and for groundingA control token only; "does not impact a site's inclusion in Google Search"
ApplebotApplePowers search features in Spotlight, Siri and SafariRespects rules for Applebot; follows Googlebot rules if Applebot isn't named
Applebot-ExtendedAppleNot a crawler: opts content out of training Apple's generative foundation models"Does not crawl webpages"; disallowing it doesn't remove pages from Apple's search results
GPTBotOpenAICrawls content that may be used to train OpenAI's generative AI foundation modelsDisallowing it indicates content should not be used in training
OAI-SearchBotOpenAISurfaces websites in ChatGPT's search featuresOpted-out sites aren't shown in ChatGPT search answers, but "can still appear as navigational links"
ChatGPT-UserOpenAIVisits pages for user actions in ChatGPT and Custom GPTs"robots.txt rules may not apply"
OAI-AdsBotOpenAIValidates the safety of pages submitted as ads on ChatGPTNot stated
ClaudeBotAnthropicCollects web content "that could potentially contribute to" training Anthropic's generative AI modelsHonours robots.txt
Claude-SearchBotAnthropicImproves search result quality for Claude usersHonours robots.txt
Claude-UserAnthropicAccesses sites when Claude users ask questionsHonours robots.txt
PerplexityBotPerplexitySurfaces and links websites in Perplexity's search results; "not used to crawl content for AI foundation models"Perplexity recommends allowing it in robots.txt
Perplexity-UserPerplexityVisits pages when users ask questions; not used for crawling or training"Generally ignores robots.txt rules"
CCBotCommon CrawlBuilds Common Crawl's open repository of web crawl dataRespects robots.txt

Search crawlers

Googlebot and Applebot crawl for search. For the AI features in Google Search, Google lists Search controls — nosnippet, data-nosnippet, max-snippet and noindex — and points to Google-Extended separately, for training and grounding in some of its other systems.

AI training crawlers

GPTBot and ClaudeBot collect content that their operators document may be used for model training. Disallowing them in robots.txt is the opt-out each operator documents, and it applies to future crawling.

AI search crawlers

OAI-SearchBot, Claude-SearchBot and PerplexityBot fetch pages so that an AI product can surface or cite them. Blocking one is a choice about appearing in that product, not about training.

User-initiated fetchers

ChatGPT-User, Claude-User and Perplexity-User fetch a page when a person asks the assistant to. Their operators document different positions on robots.txt, set out above.

Usage-control tokens

Google-Extended and Applebot-Extended are not crawlers. Google states that Google-Extended "doesn't have a separate HTTP request user agent string": crawling is done with Google's existing user agents, and the token is used "in a control capacity". Apple states that Applebot-Extended "does not crawl webpages". Neither token affects inclusion in its operator's search results.

Public content vs protected content

"Public" means served to an anonymous request. Public content can be read by any agent, including agents that ignore robots.txt, and copies of it can circulate. If something must not be read by a crawler, robots.txt is the wrong tool: put it behind authentication, or don't publish it.

What you can actually control

GoalControlWhat it doesWhat it doesn't do
Keep content privateAuthenticationRefuses every request without accessUndo copies made while it was public
Keep a page out of Google Searchnoindex (meta tag or X-Robots-Tag)Tells Google not to index the page once it is crawledWork if robots.txt blocks the page
Limit how Google Search shows a page, AI features includednosnippet, data-nosnippet, max-snippetRestricts the text Google may showRemove the page from the index
Opt out of an operator's model trainingThat operator's documented token in robots.txtRecords your choice for its future crawlingReach content already collected, or other operators
Stay out of an AI search productThat operator's search token in robots.txtAsks its search crawler not to fetchCover user-initiated fetchers that may not apply robots.txt
Reduce crawler loadrobots.txt rules for cooperating crawlersAsks them to fetch lessEnforce anything; Crawl-delay isn't part of RFC 9309, and support varies by operator
Stop unwanted trafficFirewall, CDN or rate-limit rulesRefuses the requestShow up in robots.txt, or leave every agent able to read your rules

Putting it into practice

Decide a policy per role, not per brand name

Start from the questions, not from a list of bot names: do you want to be in search indexes, in AI search products, in model training, available to assistants fetching on a user's behalf? Each answer maps to different tokens and controls. This is the part of technical and AI-search SEO that is a business decision before it is a technical one — there is no correct default for every site.

Write and place the rules

An example, not a recommendation for every site: a policy that stays open to search and AI search, and opts out of the training and dataset uses that these operators document a token for.

# Everyone else, including search and AI search crawlers: allowed
User-agent: *
Allow: /

# Training and open datasets: opted out, by each operator's documented token
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml

Consecutive User-agent lines share one group, so all five tokens receive the same rule. Two of them, Google-Extended and Applebot-Extended, are control tokens rather than crawlers: the rule records a usage preference, not a fetch restriction. CCBot builds an open dataset that others use for their own purposes, which is why it sits in this group.

To keep a single page out of search results, leave it crawlable and mark the page itself:

<meta name="robots" content="noindex">

For files without HTML, such as PDFs, send the same instruction as an HTTP response header:

X-Robots-Tag: noindex

To keep a page in Google Search but stop its text being shown as a snippet — which Google documents as also applying to its AI features — use nosnippet instead:

<meta name="robots" content="nosnippet">

Check what your rules actually say to each agent

The AI Visibility Check reads one public page and your robots.txt, and reports what your rules state for each agent in its register: allowed or disallowed, which group applied, and the operator's documented role and stance on robots.txt. It also reports the page's meta robots and X-Robots-Tag.

The register is ELIXIR's record of the operators' documentation, last checked on 27 September 2026: thirteen entries — two search crawlers, nine AI crawlers and fetchers, and two usage-control tokens. It is a reading of documentation, not an observation of crawler behaviour, and it is not exhaustive: while this article was being checked, OpenAI's page listed an agent, OAI-AdsBot, that the register does not include.

The tool states its own limit, and it applies to any reading of robots.txt: "Allowed doesn't mean a system will use or cite the page; disallowed doesn't make it invisible — and blocking at a CDN or firewall can't be seen here."

For the basics around it — whether robots.txt answers at all, the sitemap it declares, and noindex on a page — check your robots.txt, sitemap and noindex settings.

Revisit it

Tokens, purposes and policies change, and new agents appear. Re-read the operators' pages when you review your policy. Changes also take time to apply: Google generally caches robots.txt for up to 24 hours, and OpenAI says that for OAI-SearchBot it can take about 24 hours for its systems to adjust to a robots.txt update.

Sources

Checked on 1 October 2026.

THE AUDIT

Have a system that isn't working?

Describe it in the audit. ELIXIR looks at it before recommending anything.

Start an audit