KnownByLLM

Explainer · 8 min read

How big should llms.txt be?

Size, link count, and token budgets, with numbers from real files.

The spec does not say. It defines the shape of llms.txt and leaves the length to the publisher, so the practical answer has to come from what the file costs a reader and from what sites that publish one actually ship.

This explainer gives both: the sizes and link counts of real files saved in September 2026, the token arithmetic behind them, where tools draw lines, and a short set of cutting rules.

The 30-second answer: an llms.txt size limit is not in the spec, and a few kilobytes is the right order of magnitude

Aim for a file between about 2 KB and 20 KB, holding roughly 10 to 80 links with one sentence each. Half of the real files we measured are under 15 KB, and the ones that read best are in the single-digit kilobytes. Treat 50 KB as the point where you split the index into section files or move content into llms-full.txt, and 100 KB as the point where tools start to object.

The reason is cost. An assistant reads llms.txt before it reads anything else on your site, and every kilobyte of index is budget it cannot spend on your pages. A short file is read in full, so the order and annotations you wrote actually reach the model. A long one is skimmed or truncated, and the end of the file, where Optional lives, goes first.

What the spec and the tools say

  • The spec: no maximum size, no maximum link count, no guidance on either. It describes a curated index with a short sentence per link, which is a shape, not a number.
  • This site’s validator: passes a file up to 100 KiB, warns above that, and stops reading at 512 KiB, reporting a failure because the file could not be checked in full.
  • Mintlify’s generator: caps the generated index at 100,000 characters. Above that, the root llms.txt becomes a directory pointing at group files under /_llms/. A platform that generates thousands of these files decided that 100 KB is the ceiling for one index.

What real files weigh

From the 28 public files we saved on September 19, 2026, for the examples article, 24 parsed as Markdown (two fetches failed and two returned HTML). A selection, sorted by size:

SiteSizeLinksBytes per link
Svelte1.7 KB7about 240
Supabase2.7 KB32about 85
Vercel docs4.7 KB24about 200
FastHTML4.8 KB21about 230
Docker docs5.4 KB33about 160
Hono5.7 KB90about 65
Next.js12.7 KB45about 280
Cloudflare developer docs16 KB107about 150
Mintlify docs22 KB112about 195
Bun34 KB321about 105
Anthropic docs68 KB629about 110
Stripe docs92 KB453about 205
ElevenLabs docs208 KB761about 275

Two things stand out. Bytes per link cluster around 150 to 280 when every link carries a sentence, and fall to 65 to 110 when most do not: Hono and Anthropic are bare link lists. And the files above 30 KB all read as generated from a complete page inventory, while the smallest ones read as written by hand. Size is a symptom of whether anyone chose what to include.

The token arithmetic

For English prose, a working rule is about four characters per token. That makes a 5 KB file roughly 1,200 tokens, a 20 KB file about 5,000, and a 100 KB file about 25,000. Japanese and other non-Latin scripts cost more tokens per character, so a Japanese index of the same byte size is more expensive to read.

The comparison that matters is not against a model’s context window, which in 2026 is large, but against what the assistant is willing to spend on one fetched file before it has seen any of your pages. A 25,000-token index is the same cost as reading ten typical documentation pages. If the assistant has a budget for your site, you want most of it spent on the pages that answer the question, not on the list of pages.

Cutting rules

  • One sentence per link, not a paragraph. The annotation says what the page answers. Fifteen words is plenty; the files in the table average under 300 bytes per link including the URL.
  • Sections, not pages. Link the guide, not every subsection of it. If a section of your site has more than a dozen pages, link its index page and let the assistant go from there.
  • Drop what a sitemap already carries. Tag pages, pagination, legal pages, old versions, and the long tail of the blog belong in sitemap.xml, not here.
  • Split at 50 KB. Keep the root as a short directory with one link per section index, each under 20 KB. This is what Mintlify does automatically above its cap, and it works by hand too.
  • Move bodies to llms-full.txt. If the index is growing because you are pasting content into it, stop. The bundle is the place for full text; the index is the map.

One structural rule sits above the others: the first two kilobytes must stand alone. H1, blockquote, and the first section should let a reader that stops early still describe the site correctly. Everything after that is a bonus for readers with budget.

Checking the result

curl -s https://example.com/llms.txt | wc -c
curl -s https://example.com/llms.txt | grep -c '^- \['
curl -s https://example.com/llms.txt | head -c 2048

The first number is bytes, the second is links, and the third command shows what a reader with a two-kilobyte budget would see. Then run the file through a validator, which reports the size against its own threshold and lists every link so you can see where the weight is.

Validate your llms.txt

Paste your site URL and the validator fetches /llms.txt, reports its size, checks the structure against the spec, and lists every link it found.

Open the validator →

FAQ

Is there an official llms.txt size limit?

No. The llms.txt specification at llmstxt.org sets no maximum size and no maximum number of links. The limits that exist come from tools: this site's validator warns above 100 KiB and stops reading at 512 KiB, and Mintlify caps the index it generates at 100,000 characters and splits the rest into sub-index files. Treat the spec's silence as a reason to keep the file small, not as permission to make it large.

How many links is too many?

There is no number in the spec, but the files that read well in our sample have between 7 and about 110 links, with a sentence each. Above a few hundred links the file is no longer a curated index; it is a sitemap in Markdown, and an assistant gets no signal about what matters. If you need hundreds of links, group them into section files and keep the root as a short directory.

How do I estimate the token cost of my file?

For English text, a rough rule is about four characters per token, so a 10 KB file is on the order of 2,500 tokens and a 100 KB file about 25,000. Japanese and other scripts use more tokens per character, so the same byte count costs more. The number to compare against is not a model's context window but the share of it an assistant is willing to spend on one fetched file before it has read a single page of yours.

Should I move content to llms-full.txt instead of growing llms.txt?

Yes, when what is growing is page content rather than the index. llms.txt is the table of contents and should stay small enough to read in full; llms-full.txt is the bundle of full-text pages and can be hundreds of kilobytes. Growing the index with long descriptions or inline content is the common mistake; the fix is a short annotation per link and a separate bundle.

Does file size affect whether AI tools read llms.txt at all?

Not that anyone has documented. Whether a given assistant fetches the file depends on the vendor, not on its size. What size affects is what happens after the fetch: a short file is read whole and its priorities are visible, while a very long one is more likely to be truncated or skimmed, and the sections at the end, including Optional, are the first to be lost.

Next steps