KnownByLLM

Reference · 11 min read

AI crawlers list: which bots to allow in robots.txt

Every documented AI user agent as of September 2026, sorted by what it actually does, with three robots.txt recipes.

Most “block the AI bots” lists treat every user agent the same. They are not the same. Some feed training sets, some build search indexes, some fetch a single page because a person asked, and two of them are not crawlers at all. Blocking the wrong one removes you from AI answers while leaving your content in training data, which is the opposite of what most site owners intend.

This list was compiled from each vendor’s own crawler documentation on 24 September 2026. Where a vendor has no documentation, the article says so rather than guessing.

The 30-second answer: AI crawlers by job

Sort every AI user agent into one of three jobs and the robots.txt decision becomes simple.

  • Training crawlers collect content that may be used to train models: GPTBot, ClaudeBot, meta-externalagent, CCBot, Amazonbot, plus the two tokens Google-Extended and Applebot-Extended. Blocking these does not affect whether you appear in that vendor’s search or answers.
  • Search indexers build the index behind AI search features: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer, Amzn-SearchBot, and the ordinary Googlebot and Applebot. Block these and you disappear from those results.
  • User-initiated fetchers load a page at answer time because a person asked: ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Amzn-User. Most vendors document that these may ignore robots.txt.

The common mistake is a single Disallow: / for everything with “bot” in the name. If your goal is “do not train on my content but do cite me”, block the first group and allow the other two. The recipes below do exactly that.

The list, by vendor

Purposes are paraphrased from the vendor’s documentation; the robots.txt column is what the vendor states, not what logs prove.

User agentVendorJobrobots.txt
GPTBotOpenAITrainingRespected
OAI-SearchBotOpenAISearch index (ChatGPT search)Respected; must be allowed to appear
ChatGPT-UserOpenAIUser fetchMay not apply
ClaudeBotAnthropicTrainingRespected, plus Crawl-delay
Claude-SearchBotAnthropicSearch indexRespected
Claude-UserAnthropicUser fetchRespected (stated)
GooglebotGoogleSearch index (incl. AI Overviews)Respected
Google-ExtendedGoogleTraining and grounding token; not a crawlerToken only; no effect on Search
PerplexityBotPerplexitySearch indexRespected
Perplexity-UserPerplexityUser fetchGenerally ignored (stated)
meta-externalagentMetaTraining or indexingRespected
Meta-WebIndexerMetaSearch index (Meta AI)Respected
Meta-ExternalFetcherMetaUser fetchMay bypass
ApplebotAppleSearch index (Siri, Spotlight)Respected
Applebot-ExtendedAppleTraining token; not a crawlerToken only; no effect on search
AmazonbotAmazonProducts, may train modelsRespected; cached up to 30 days
Amzn-SearchBot / Amzn-UserAmazonSearch / user fetch; no trainingRespected
CCBotCommon CrawlOpen crawl corpus (used by many trainers)Respected
BytespiderByteDanceUndocumentedNo vendor documentation found

Notes on the table

  • OpenAI also lists OAI-AdsBot, which only checks pages submitted as ChatGPT ads. It is omitted above as irrelevant to most sites.
  • Anthropic is the one vendor whose documentation (updated April 2026) says its user-initiated fetcher, Claude-User, respects robots.txt, and that all three bots honour Crawl-delay.
  • Meta additionally runs facebookexternalhit for link previews when someone shares your page, and Meta-ExternalAds for advertising. Neither is an AI crawler in the sense above.
  • Bytespider is widely reported as a ByteDance training crawler, but we could find no vendor page describing its purpose, its robots.txt behaviour or its IP ranges. Treat it as an unknown: block it by user agent if you wish, and do not rely on that alone.

Two tokens that will never show up in your logs

Google-Extended and Applebot-Extended cause more confusion than any real crawler. Neither makes requests. Google’s crawler page says Google-Extended controls whether content may be used to train future Gemini models and for grounding in Gemini apps and Vertex AI, and that it does not impact inclusion in Search and is not a ranking signal. Apple’s page says Applebot-Extended does not crawl webpages, that pages disallowing it can still appear in search results, and that its rules are not considered in ranking.

So: you set rules for these tokens in robots.txt, the real crawler (Googlebot, Applebot) reads those rules and applies them to how the fetched content may be used. Searching your access log for either name will always return nothing, and that is not a sign anything is wrong.

Three robots.txt recipes

Recipe 1: allow everything (the default)

If your robots.txt has no rules for these agents, they are allowed. You do not need to add Allow lines. The only reason to write them out is to override a broader block, for example a WAF or a template that disallows unknown bots.

Recipe 2: cite me, but do not train on me

The most common intent. Block the training group, say nothing about the search and user-fetch group.

# Training crawlers and tokens: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: CCBot
User-agent: Bytespider
Disallow: /

# Search indexers and user fetchers stay allowed by default:
# OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User,
# PerplexityBot, Perplexity-User, Googlebot, Applebot ...

User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml

Grouping several User-agent lines above one Disallow is valid in the Robots Exclusion Protocol and is what Google’s own parser expects. Note that Amazon says Amazonbot may cache your robots.txt for up to 30 days, so a change there can take a month to bite.

Recipe 3: no AI at all

Block all three groups. Understand what you are giving up: no citations in ChatGPT, Claude or Perplexity answers, and, because Meta and Amazon bundle search and AI, reduced presence in those products too. Do not block Googlebot or Applebot here unless you also want to leave ordinary web search.

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Google-Extended
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-externalagent
User-agent: Meta-WebIndexer
User-agent: Meta-ExternalFetcher
User-agent: Applebot-Extended
User-agent: Amazonbot
User-agent: CCBot
User-agent: Bytespider
Disallow: /

Even here, the vendors that document user-initiated fetchers (OpenAI, Perplexity, Meta) say those may still load a page a person explicitly asks for. robots.txt is a request, not a firewall. If you need enforcement, that is a WAF or bot management setting, not a text file.

Verifying that a bot is who it says it is

A user-agent string costs nothing to fake, and a large share of “GPTBot” traffic in the wild is scrapers borrowing the name. Before acting on a log line, check the source IP against the vendor’s published ranges:

  • OpenAI: openai.com/gptbot.json, searchbot.json, chatgpt-user.json
  • Anthropic: claude.com/crawling/bots.json
  • Perplexity: perplexity.com/perplexitybot.json and perplexity-user.json
  • Apple: search.developer.apple.com/applebot.json, or reverse DNS to *.applebot.apple.com
  • Amazon: the IP address pages linked from developer.amazon.com/amazonbot
  • Common Crawl: index.commoncrawl.org/ccbot.json, or reverse DNS to *.crawl.commoncrawl.org
  • Google: the verification method on Google’s crawler documentation (reverse DNS to googlebot.com or google.com)

A request claiming to be GPTBot from an IP outside OpenAI’s list is not GPTBot, and robots.txt rules for GPTBot will not change its behaviour.

If you use Cloudflare

Cloudflare offers a managed robots.txt setting (documentation updated August 2026) that prepends rules to your file. As of that page it adds a Content-Signal line (search allowed, AI training disallowed) and a Disallow: / for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. That is close to Recipe 2, it is opt-in, and it is available on all plans including Free. Check what it generates before relying on it; the list of bots it targets can change, and it does not verify IPs for you.

Where llms.txt fits

Nowhere in the access decision. llms.txt describes your site to an agent that is already allowed in; robots.txt decides who is allowed in. The llms.txt spec states the two files have different purposes. The one practical interaction: a crawler you disallow in robots.txt will not fetch your llms.txt either, so Recipe 3 makes the file invisible to those vendors.

See which bots can reach your site

Run the checker on your URL. It fetches your robots.txt and reports what the common AI user agents are allowed to read, alongside the llms.txt draft it generates.

Run the check →

FAQ

Which AI crawlers should I allow in robots.txt?

If you want to appear in AI answers, allow the search and user-fetch bots: OAI-SearchBot and ChatGPT-User (OpenAI), Claude-SearchBot and Claude-User (Anthropic), PerplexityBot and Perplexity-User, and Googlebot as usual. Whether to allow the training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, Amazonbot) is a separate decision that does not affect search visibility, according to each vendor's documentation.

If I block GPTBot, will my site disappear from ChatGPT?

Not from ChatGPT search. OpenAI documents GPTBot as the crawler for content that may be used in training, and OAI-SearchBot as the one that surfaces sites in ChatGPT search. Blocking GPTBot alone leaves OAI-SearchBot and ChatGPT-User unaffected. The same split applies at Anthropic (ClaudeBot vs Claude-SearchBot) and Apple (Applebot-Extended vs Applebot).

Why do Google-Extended and Applebot-Extended never appear in my logs?

Because they are not crawlers. They are robots.txt tokens that the real crawlers (Googlebot, Applebot) read to decide whether fetched content may be used for AI training. Google and Apple both state this explicitly, and both say the token does not affect search inclusion or ranking. You cannot grep for them; you can only set rules for them.

Do AI crawlers actually obey robots.txt?

The documented ones say they do, with one deliberate exception: user-initiated fetchers. OpenAI's ChatGPT-User, Perplexity's Perplexity-User and Meta's Meta-ExternalFetcher are documented as able to bypass robots.txt because a person asked for the page. Anthropic, unusually, states that Claude-User does respect robots.txt. Undocumented crawlers cannot be assumed to obey anything.

Does llms.txt have anything to do with this?

No. llms.txt does not control access; robots.txt does. The llms.txt spec says so itself. Blocking a crawler in robots.txt also stops it fetching your llms.txt, so the two files interact only in that direction.

How do I verify that a request really came from OpenAI or Anthropic?

Do not trust the user-agent string; it is trivially faked. Each major vendor publishes IP ranges: openai.com/gptbot.json (and searchbot.json, chatgpt-user.json), claude.com/crawling/bots.json, perplexity.com/perplexitybot.json, search.developer.apple.com/applebot.json, Amazon's amazonbot IP pages, and index.commoncrawl.org/ccbot.json. Compare the request IP against the list, or use reverse DNS where the vendor documents it.

Next steps