The 30-second answer: AI crawlers by job
Sort every AI user agent into one of three jobs and the robots.txt decision becomes simple.
- Training crawlers collect content that may be used to train models: GPTBot, ClaudeBot, meta-externalagent, CCBot, Amazonbot, plus the two tokens Google-Extended and Applebot-Extended. Blocking these does not affect whether you appear in that vendor’s search or answers.
- Search indexers build the index behind AI search features: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer, Amzn-SearchBot, and the ordinary Googlebot and Applebot. Block these and you disappear from those results.
- User-initiated fetchers load a page at answer time because a person asked: ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Amzn-User. Most vendors document that these may ignore robots.txt.
The common mistake is a single Disallow: / for everything with “bot” in the name. If your goal is “do not train on my content but do cite me”, block the first group and allow the other two. The recipes below do exactly that.
The list, by vendor
Purposes are paraphrased from the vendor’s documentation; the robots.txt column is what the vendor states, not what logs prove.
| User agent | Vendor | Job | robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Training | Respected |
| OAI-SearchBot | OpenAI | Search index (ChatGPT search) | Respected; must be allowed to appear |
| ChatGPT-User | OpenAI | User fetch | May not apply |
| ClaudeBot | Anthropic | Training | Respected, plus Crawl-delay |
| Claude-SearchBot | Anthropic | Search index | Respected |
| Claude-User | Anthropic | User fetch | Respected (stated) |
| Googlebot | Search index (incl. AI Overviews) | Respected | |
| Google-Extended | Training and grounding token; not a crawler | Token only; no effect on Search | |
| PerplexityBot | Perplexity | Search index | Respected |
| Perplexity-User | Perplexity | User fetch | Generally ignored (stated) |
| meta-externalagent | Meta | Training or indexing | Respected |
| Meta-WebIndexer | Meta | Search index (Meta AI) | Respected |
| Meta-ExternalFetcher | Meta | User fetch | May bypass |
| Applebot | Apple | Search index (Siri, Spotlight) | Respected |
| Applebot-Extended | Apple | Training token; not a crawler | Token only; no effect on search |
| Amazonbot | Amazon | Products, may train models | Respected; cached up to 30 days |
| Amzn-SearchBot / Amzn-User | Amazon | Search / user fetch; no training | Respected |
| CCBot | Common Crawl | Open crawl corpus (used by many trainers) | Respected |
| Bytespider | ByteDance | Undocumented | No vendor documentation found |
Notes on the table
- OpenAI also lists OAI-AdsBot, which only checks pages submitted as ChatGPT ads. It is omitted above as irrelevant to most sites.
- Anthropic is the one vendor whose documentation (updated April 2026) says its user-initiated fetcher, Claude-User, respects robots.txt, and that all three bots honour
Crawl-delay. - Meta additionally runs facebookexternalhit for link previews when someone shares your page, and Meta-ExternalAds for advertising. Neither is an AI crawler in the sense above.
- Bytespider is widely reported as a ByteDance training crawler, but we could find no vendor page describing its purpose, its robots.txt behaviour or its IP ranges. Treat it as an unknown: block it by user agent if you wish, and do not rely on that alone.
Two tokens that will never show up in your logs
Google-Extended and Applebot-Extended cause more confusion than any real crawler. Neither makes requests. Google’s crawler page says Google-Extended controls whether content may be used to train future Gemini models and for grounding in Gemini apps and Vertex AI, and that it does not impact inclusion in Search and is not a ranking signal. Apple’s page says Applebot-Extended does not crawl webpages, that pages disallowing it can still appear in search results, and that its rules are not considered in ranking.
So: you set rules for these tokens in robots.txt, the real crawler (Googlebot, Applebot) reads those rules and applies them to how the fetched content may be used. Searching your access log for either name will always return nothing, and that is not a sign anything is wrong.
Three robots.txt recipes
Recipe 1: allow everything (the default)
If your robots.txt has no rules for these agents, they are allowed. You do not need to add Allow lines. The only reason to write them out is to override a broader block, for example a WAF or a template that disallows unknown bots.
Recipe 2: cite me, but do not train on me
The most common intent. Block the training group, say nothing about the search and user-fetch group.
# Training crawlers and tokens: opt out User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: meta-externalagent User-agent: Amazonbot User-agent: CCBot User-agent: Bytespider Disallow: / # Search indexers and user fetchers stay allowed by default: # OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, # PerplexityBot, Perplexity-User, Googlebot, Applebot ... User-agent: * Allow: / Sitemap: https://example.com/sitemap.xml
Grouping several User-agent lines above one Disallow is valid in the Robots Exclusion Protocol and is what Google’s own parser expects. Note that Amazon says Amazonbot may cache your robots.txt for up to 30 days, so a change there can take a month to bite.
Recipe 3: no AI at all
Block all three groups. Understand what you are giving up: no citations in ChatGPT, Claude or Perplexity answers, and, because Meta and Amazon bundle search and AI, reduced presence in those products too. Do not block Googlebot or Applebot here unless you also want to leave ordinary web search.
User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: Google-Extended User-agent: PerplexityBot User-agent: Perplexity-User User-agent: meta-externalagent User-agent: Meta-WebIndexer User-agent: Meta-ExternalFetcher User-agent: Applebot-Extended User-agent: Amazonbot User-agent: CCBot User-agent: Bytespider Disallow: /
Even here, the vendors that document user-initiated fetchers (OpenAI, Perplexity, Meta) say those may still load a page a person explicitly asks for. robots.txt is a request, not a firewall. If you need enforcement, that is a WAF or bot management setting, not a text file.
Verifying that a bot is who it says it is
A user-agent string costs nothing to fake, and a large share of “GPTBot” traffic in the wild is scrapers borrowing the name. Before acting on a log line, check the source IP against the vendor’s published ranges:
- OpenAI:
openai.com/gptbot.json,searchbot.json,chatgpt-user.json - Anthropic:
claude.com/crawling/bots.json - Perplexity:
perplexity.com/perplexitybot.jsonandperplexity-user.json - Apple:
search.developer.apple.com/applebot.json, or reverse DNS to*.applebot.apple.com - Amazon: the IP address pages linked from
developer.amazon.com/amazonbot - Common Crawl:
index.commoncrawl.org/ccbot.json, or reverse DNS to*.crawl.commoncrawl.org - Google: the verification method on Google’s crawler documentation (reverse DNS to
googlebot.comorgoogle.com)
A request claiming to be GPTBot from an IP outside OpenAI’s list is not GPTBot, and robots.txt rules for GPTBot will not change its behaviour.
If you use Cloudflare
Cloudflare offers a managed robots.txt setting (documentation updated August 2026) that prepends rules to your file. As of that page it adds a Content-Signal line (search allowed, AI training disallowed) and a Disallow: / for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. That is close to Recipe 2, it is opt-in, and it is available on all plans including Free. Check what it generates before relying on it; the list of bots it targets can change, and it does not verify IPs for you.
Where llms.txt fits
Nowhere in the access decision. llms.txt describes your site to an agent that is already allowed in; robots.txt decides who is allowed in. The llms.txt spec states the two files have different purposes. The one practical interaction: a crawler you disallow in robots.txt will not fetch your llms.txt either, so Recipe 3 makes the file invisible to those vendors.
See which bots can reach your site
Run the checker on your URL. It fetches your robots.txt and reports what the common AI user agents are allowed to read, alongside the llms.txt draft it generates.
Run the check →FAQ
Which AI crawlers should I allow in robots.txt?
If you want to appear in AI answers, allow the search and user-fetch bots: OAI-SearchBot and ChatGPT-User (OpenAI), Claude-SearchBot and Claude-User (Anthropic), PerplexityBot and Perplexity-User, and Googlebot as usual. Whether to allow the training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, Amazonbot) is a separate decision that does not affect search visibility, according to each vendor's documentation.
If I block GPTBot, will my site disappear from ChatGPT?
Not from ChatGPT search. OpenAI documents GPTBot as the crawler for content that may be used in training, and OAI-SearchBot as the one that surfaces sites in ChatGPT search. Blocking GPTBot alone leaves OAI-SearchBot and ChatGPT-User unaffected. The same split applies at Anthropic (ClaudeBot vs Claude-SearchBot) and Apple (Applebot-Extended vs Applebot).
Why do Google-Extended and Applebot-Extended never appear in my logs?
Because they are not crawlers. They are robots.txt tokens that the real crawlers (Googlebot, Applebot) read to decide whether fetched content may be used for AI training. Google and Apple both state this explicitly, and both say the token does not affect search inclusion or ranking. You cannot grep for them; you can only set rules for them.
Do AI crawlers actually obey robots.txt?
The documented ones say they do, with one deliberate exception: user-initiated fetchers. OpenAI's ChatGPT-User, Perplexity's Perplexity-User and Meta's Meta-ExternalFetcher are documented as able to bypass robots.txt because a person asked for the page. Anthropic, unusually, states that Claude-User does respect robots.txt. Undocumented crawlers cannot be assumed to obey anything.
Does llms.txt have anything to do with this?
No. llms.txt does not control access; robots.txt does. The llms.txt spec says so itself. Blocking a crawler in robots.txt also stops it fetching your llms.txt, so the two files interact only in that direction.
How do I verify that a request really came from OpenAI or Anthropic?
Do not trust the user-agent string; it is trivially faked. Each major vendor publishes IP ranges: openai.com/gptbot.json (and searchbot.json, chatgpt-user.json), claude.com/crawling/bots.json, perplexity.com/perplexitybot.json, search.developer.apple.com/applebot.json, Amazon's amazonbot IP pages, and index.commoncrawl.org/ccbot.json. Compare the request IP against the list, or use reverse DNS where the vendor documents it.
Next steps
- → llms.txt vs robots.txt vs sitemap.xml (how the three root files divide the work)
- → Does ChatGPT actually use llms.txt? (what OpenAI’s bots do and do not fetch)
- → All Learn articles