The 30-second answer: an AI crawler vs Googlebot is not one comparison but three, and on every axis the AI side is simpler
Googlebot is one crawler that feeds one index, renders JavaScript with an evergreen Chromium in a queued second phase, sets its own crawl rate, and reads robots.txt by published rules. The AI side is three different jobs, each run by several vendors: training crawlers that collect text, search-index crawlers that feed an answer engine, and user fetchers that load one page because someone asked about it.
None of the vendors documents JavaScript rendering, so assume they read the HTML as served. The user fetchers have no schedule and no index; they read the page at answer time, and two of them say robots.txt may not apply because a person initiated the request. Everything else about serving them follows from those two facts.
Three jobs, not one crawler
The vendor documentation separates the agents by purpose, and the separation matters more than the vendor:
- Training crawlers (GPTBot, ClaudeBot) collect content that may go into a model. OpenAI says disallowing GPTBot indicates a site’s content should not be used in training; Anthropic describes ClaudeBot as collecting web content that could contribute to training.
- Search-index crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) build the index an answer engine cites from. OpenAI says a site must allow OAI-SearchBot to appear in ChatGPT search; Perplexity says PerplexityBot is not used to crawl content for foundation models.
- User fetchers (ChatGPT-User, Claude-User, Perplexity-User) load a page when a person asks about it. OpenAI states ChatGPT-User is not used for crawling the web in an automatic fashion.
Google has one crawler for all of Search including AI Overviews, and a control token, Google-Extended, which has no user agent of its own and, per Google, does not affect inclusion in Search or ranking. The full list with vendors is in the crawler list article.
The comparison
| Googlebot | AI training crawlers | AI search crawlers | AI user fetchers | |
|---|---|---|---|---|
| Feeds | The Search index, including AI Overviews and AI Mode | A training corpus | An answer engine's index | One answer, once |
| JavaScript | Rendered with an evergreen Chromium, queued after crawling | Not documented; assume HTML only | Not documented; assume HTML only | Not documented; assume HTML only |
| When it reads a page | On its own schedule, before any query | Periodically, before any query | Periodically, before any query | At answer time, because a user asked |
| Keeps a copy | Yes, the index | Yes, in training data | Yes, the index | No schedule, no index documented |
| Crawl rate | Set by Google's algorithms; 429/500/503 reduce it host-wide | Vendor-controlled; Anthropic honors Crawl-delay | Vendor-controlled | One request per question |
| robots.txt | Documented rules: most specific group, longest path, least restrictive on tie | Respected (vendor statement) | Respected; must be allowed to be cited | May not apply (OpenAI); generally ignored (Perplexity); respected (Anthropic) |
| Blocking costs you | Search visibility | Nothing visible | Citations in that engine | The answer a user gets right now |
JavaScript: one renderer, and silence everywhere else
Google processes JavaScript in three phases, crawling, rendering, and indexing, and says Search runs JavaScript with an evergreen version of Chromium. The rendering happens in a queue; Google says a page may stay on it for a few seconds but it can take longer. That is why a client-rendered page can be indexed with its full content, late.
The OpenAI, Anthropic, and Perplexity bot pages do not mention rendering at all. What server logs show is an HTTP request for the HTML and nothing more: no follow-up requests for script bundles or JSON the page would load in a browser. The safe assumption is that every AI agent reads the HTML response as served and stops. Content that exists only after a script runs, behind a click, or in a fetch the page makes after load does not exist for them.
The practical consequence is the oldest advice in SEO, now with a sharper reason: server-render the content that matters. A pricing table, a product description, a documentation page should be in the initial HTML. Where that is hard, the llms.txt spec’s suggestion applies: provide a clean Markdown version of the page at the same URL with .md appended, and point to it from llms.txt.
Timing and frequency: a schedule versus a question
Googlebot decides its own rate. Google says its infrastructure uses algorithms to determine the optimal crawl rate for a site, that a significant number of 429, 500, or 503 responses reduces the crawl rate across the whole hostname, and that the rate increases again when the errors stop. Google also warns that a URL returning those codes for multiple days may be dropped from the index. The index is the product; crawling is how it stays current.
Training and search crawlers from AI vendors also work ahead of any query, but none publishes a frequency. What is documented is the cadence of control changes: Google caches robots.txt for up to 24 hours; Perplexity says changes may take up to 24 hours to be reflected; Amazon says Amazonbot may cache robots.txt for up to 30 days. A robots.txt change is not instant for anyone.
The user fetchers have no schedule. ChatGPT-User, Claude-User, and Perplexity-User arrive when a person asks a question that needs your page, read it, and answer. There is no documented index behind them, which is why a page changed this morning can be quoted correctly this afternoon, and why a page they cannot read is simply absent from the answer. The llms.txt spec describes the same pattern for its file: information used on demand, when an agent needs it.
robots.txt: documented rules versus vendor statements
Google documents how it reads the file. Only one group applies to a crawler: the one with the most specific user-agent match. Within the group, the most specific rule by path length wins, and when an allow and a disallow of equal specificity conflict, Google uses the least restrictive rule. A 4xx response other than 429 is treated as if no robots.txt existed; a 5xx stops crawling for the first 12 hours while Google retries. Google does not support Crawl-delay.
The AI vendors state intent rather than parsing rules. OpenAI and Anthropic say their crawlers respect robots.txt; Anthropic adds that it honors the non-standard Crawl-delay extension. Perplexity asks to be allowed. For the user fetchers the statements diverge: OpenAI says robots.txt rules may not apply to ChatGPT-User because a user initiated the action; Perplexity says Perplexity-User generally ignores robots.txt for the same reason; Anthropic says Claude-User respects it. Two tokens, Google-Extended and Applebot-Extended, are not crawlers at all: they control use of what the ordinary crawler already fetched.
Two consequences for the file you write. Name groups explicitly, because a wildcard group is what an AI crawler falls back to when it finds no match, and because Google will pick the most specific group and ignore the rest:
User-agent: * Allow: / User-agent: GPTBot User-agent: ClaudeBot Disallow: / User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot Allow: /
And accept that robots.txt is not a lock for user fetchers. If a page must not be read by an assistant on a user’s behalf, the only reliable control is the server: authentication, or a firewall rule, with the cost that the user gets no answer from your site.
What this means for serving a site
- Put the content in the HTML. Googlebot will eventually render your scripts; nothing else documents that it will. Server-side rendering, static generation, or a Markdown twin of the page covers all readers at once.
- Decide per job, not per vendor. Blocking a training crawler costs nothing visible. Blocking a search crawler removes you from that engine’s citations. Blocking a user fetcher removes you from one person’s answer, and may not work.
- Do not rate-limit by user-agent string.Anything can claim to be Googlebot or GPTBot. Verify by published IP ranges or reverse DNS, and rate-limit by address and behavior.
- Expect robots changes to lag. Up to a day for Google and Perplexity by their own statements, up to a month for Amazonbot. Make the change, then wait before judging it.
- Give the on-demand readers an index. A user fetcher has one request and a few thousand tokens. An llms.txt at the root tells it which page to read for which question; the writing guide takes about 30 minutes.
Checking who can reach you
Fetch your robots.txt and a key page with the exact user-agent strings above, from outside your network, and compare what comes back with what a browser shows. A checker does the same for the common agents in one pass.
See which bots can reach your site
Run the checker on your URL. It fetches your robots.txt and reports what the common AI user agents are allowed to read, alongside the llms.txt draft it generates.
Run the check →FAQ
Do AI crawlers render JavaScript like Googlebot?
Google documents that Search runs JavaScript with an evergreen version of Chromium, in a rendering phase that is queued after crawling. The OpenAI, Anthropic, and Perplexity bot documentation says nothing about rendering, and the fetches seen in server logs are plain HTTP requests for the HTML. Treat AI crawlers and fetchers as readers of the HTML response only: content that appears after script execution should not be assumed visible to them.
Is an AI crawler one thing?
No. Each vendor runs up to three agents with different jobs: a training crawler (GPTBot, ClaudeBot), a search-index crawler (OAI-SearchBot, Claude-SearchBot, PerplexityBot), and a user-initiated fetcher (ChatGPT-User, Claude-User, Perplexity-User). They differ in timing, in how they treat robots.txt, and in what blocking them costs. Googlebot is one crawler feeding one index; the comparison only makes sense per job.
Do AI bots respect robots.txt?
The training and search crawlers say they do. The user-initiated fetchers are the exception: OpenAI says robots.txt rules may not apply to ChatGPT-User because a user initiated the request, and Perplexity says Perplexity-User generally ignores robots.txt for the same reason. Anthropic states its bots respect robots.txt. Google-Extended is a control token, not a crawler, and has no effect on Search.
How often do they come back?
Googlebot sets its own rate with algorithms that react to your server: many 429, 500, or 503 responses reduce the crawl rate across the whole hostname, and it recovers when they stop. Vendor documentation for AI crawlers does not publish a frequency. The user fetchers have no schedule at all: they arrive when someone asks, read the page once, and the spec for llms.txt describes the same on-demand pattern.
Does Crawl-delay work?
Not for Google, which does not support the rule. Anthropic documents support for the non-standard Crawl-delay extension and says it respects it where appropriate. For other vendors, treat it as unsupported and use server-side rate limiting instead, which works regardless of what the bot claims to honor.
How do I tell a real Googlebot or GPTBot from an impostor?
Verify by network, not by user-agent string. Google publishes its crawler IP ranges and supports reverse DNS lookup; OpenAI publishes IP ranges for its bots. A request that claims to be GPTBot from an address outside the published ranges is not GPTBot. The crawler list article on this site has the verification steps.
Next steps
- → AI crawlers list: which bots to allow in robots.txt (every user agent by vendor, with verification steps)
- → Does Google use llms.txt? (what governs AI Overviews and AI Mode)
- → Does ChatGPT actually use llms.txt?
- → All Learn articles