KnownByLLM

Reference · 10 min read

The llms.txt spec, explained line by line

What v2 actually requires, what it leaves open, and where readers get it wrong.

The llms.txt specification is short: a few paragraphs of motivation, a bulleted definition of the format, one example and a handful of guidelines. That brevity is why so much folklore has grown around it. This article goes through the normative parts one clause at a time and says, for each, whether it is required, optional or simply not addressed.

It is based on the text of llmstxt.org as fetched on 23 September 2026: version 2 of the proposal, revised in August 2026. Where v2 changed something from the 2024 original, that is called out.

The 30-second answer: what the llms.txt spec requires

A conforming file is a Markdown document named llms.txt, served at the site root or at any sub-path, containing in this order: an optional byte-order mark, an H1 with the site or project name, an optional blockquote summary, optional free-form Markdown without headings, and zero or more H2 sections each holding a Markdown list whose items are a link, optionally followed by a colon and notes. The H1 is the only required element.

Everything else people argue about, from Content-Type to maximum size to whether llms-full.txt is mandatory, is outside the spec. The rest of this article is the evidence for that sentence.

Where the file lives

The spec: files named llms.txt, at the root path or at any sub-path such as /docs/llms.txt. A file covers the URLs under its path, and where more than one file applies, agents should use the most specific one.

This is a v2 change. The 2024 text allowed sub-paths without saying what they meant; v2 defines it, which is what lets a site that only controls one path, such as a project site on shared hosting, participate. The spec also explains why it did not use /.well-known/: well-known URIs exist only at the origin root, and an llms.txt is meant to describe the path where it sits, like index.html does.

The format, clause by clause

The definition is a single bulleted list. Here it is, with what each bullet does and does not mean.

“An optional byte-order mark”

A UTF-8 BOM at the start of the file is permitted. That is the only thing the spec says about encoding: it implies UTF-8 without naming it, and says nothing about Content-Type. Serve it as text/plain or text/markdown with a UTF-8 charset and you match every large adopter.

“An H1 with the name of the project or site. This is the only required section”

One line starting with #. Required, and the only required thing. The spec does not say it must be the first line (the BOM may precede it) or that there must be exactly one, but the intent is clearly a single title, and validators treat a second H1 as at least a warning.

“A blockquote with a short summary of the project”

A paragraph starting with >, containing, in the spec’s words, key information necessary for understanding the rest of the file. Optional in the strict sense, and the spec’s own mock example labels it “Optional description”. It is the single most useful line for an agent, so leaving it out is legal and unwise. Two of the eight large files we examined last week omit it.

“Zero or more markdown sections … of any type except headings”

Between the summary and the first H2, you may put paragraphs, lists, even code blocks, containing more detailed information about the project and how to interpret the provided files. The one rule is no headings, because a heading would start a file list. This is where FastHTML puts its “things to remember” and Stripe its instructions for agents.

“Zero or more markdown sections delimited by H2 headers, containing file lists”

Each ## heading names a section and is followed by a Markdown list. The spec says H2, not H3; nested H3 groups inside a section, as Anthropic’s file uses, are not described, though parsers that only look for H2 boundaries and list items handle them fine. Section names are free text. Only “Optional” carries a conventional meaning.

“A required markdown hyperlink, then optionally a colon and notes”

The list item syntax is - [name](url): notes. The link is required; the colon and notes are optional. The spec says “markdown list” without fixing the bullet character, so * is valid Markdown, but - is what the example uses and what every strict parser expects. URLs are not restricted to your own host; the spec explicitly counts links to external sites as an advantage over sitemaps.

The Optional section and what v2 changed

The spec: the Optional section is used, by convention, for secondary information, links an agent can skip when a shorter context is needed.

In v1 this had teeth. The original proposal described a context-expansion tool that would include everything except Optional when building a prompt. v2 drops that tooling from the proposal and, with it, the special meaning: Optional sections are still allowed and remain a useful convention, but they no longer carry mechanical semantics. If a guide tells you agents will “always skip” Optional, it is describing v1.

What the spec asks of the linked pages

Two recommendations sit outside the file itself and are easy to miss.

  • Markdown versions of pages. Pages that agents might need should offer a clean Markdown version at the same URL with .md appended (page.html.md) or the extension replaced (page.md); a directory URL uses index.html.md or index.md. v1 specified only the appended form; v2 allows both because tooling had diverged.
  • Link relations. v2 adds a discovery mechanism: rel="alternate" type="text/markdown" points from a page to its Markdown version, and rel="describedby" points to the llms.txt that covers it. Either can be an HTML <link> element or an HTTP Link: header, and the header form can be added at the CDN without touching pages.
Link: </docs/page.html.md>; rel="alternate"; type="text/markdown", </docs/llms.txt>; rel="describedby"

Neither is required for a valid llms.txt. Both are what the spec means when it says links should point to LLM-friendly content. If your site cannot produce Markdown versions, linking to HTML pages is still conforming.

What the spec does not say

The gaps matter as much as the rules, because most of the confident advice online lives in them.

TopicWhat the spec saysWhat people assume
Content-TypeNothingThat text/plain is mandatory; text/markdown is equally common
File size or link countNothing numeric; only that the file should fit in contextHard limits such as 25 links or 10 KB
llms-full.txtNot mentionedThat it is part of the standard or required
Which bots read itAgents view or search it, then follow links; no consumers are namedThat ChatGPT or Google Search index it
Bullet character and H3A markdown list; H2 sectionsThat * bullets or H3 sub-sections are invalid
Training vs inferenceExpected mainly for inference, though training could use itThat it grants or denies training permission

The last row is the one with consequences. The spec is explicit that robots.txt and llms.txt have different purposes: robots.txt says what access is acceptable, llms.txt is read on demand when an agent needs information. Nothing in llms.txt controls crawling or training.

The four guidelines

After the example, the spec offers four sentences of advice. They are not normative, but they are the best summary of what makes a file useful:

  • Use concise, clear language.
  • When linking to resources, include brief, informative descriptions.
  • Avoid ambiguous terms or unexplained jargon.
  • Test your file by asking an agent questions about your content, giving it only your llms.txt as a starting point.

The fourth is the one nobody does. It is also the only test that measures what the file is for.

Check a file against these rules

The validator applies the normative parts above: H1 present and single, summary present, H2 sections with well-formed link items, absolute URLs, plus the practical checks the spec leaves open such as encoding and size.

Open the validator →

A minimal conforming file

Everything required, and the two optional parts every good file includes:

# Example Co

> Example Co makes widgets for small workshops in Ireland, since 2012.

## Pages

- [Products](https://example.com/products.md): The three widget models, with prices and lead times.
- [Contact](https://example.com/contact.md): Phone, email and the booking form for a site visit.

Strip the blockquote and the section and you are left with # Example Co, which is valid and pointless. The spec sets a low floor on purpose; the guidelines and the real files set the standard.

FAQ

Where is the official llms.txt spec?

At https://llmstxt.org/, maintained by Jeremy Howard of Answer.AI, with the source in the AnswerDotAI/llms-txt GitHub repository. The current text is v2, revised in August 2026; the original proposal dates from September 2024. The page also lists what changed between v1 and v2.

What is the only required part of an llms.txt file?

The H1 line with the name of the site or project. The spec says so explicitly. The blockquote summary, the free-form notes and the H2 link sections are all defined but optional. In practice a file with only an H1 is useless to an agent, so treat the summary and at least one link section as required for a good file, not for a valid one.

Does llms.txt have to be at the site root?

Not any more. v2 says the file can live at the root or at any sub-path, and that a file covers the URLs under its path; when several apply, an agent should use the most specific one. /docs/llms.txt is the canonical example. For a small site the root file is still the normal choice.

Is the Optional section special?

Only by convention. In v1 it told a context-expansion tool which links to drop. v2 removed that tooling and with it the mechanical meaning, but keeps Optional as the conventional place for links an agent can skip when context is short.

What does the spec say about Content-Type, encoding or file size?

Nothing, beyond allowing an optional byte-order mark at the start. Content-Type, charset and size are left to the publisher. text/plain or text/markdown with UTF-8 is what the large adopters use, and keeping the file small is implied by the design goal that it fit in an agent's context.

Is llms-full.txt part of the spec?

No. The spec defines llms.txt and recommends Markdown versions of individual pages; it does not define llms-full.txt. That file is a separate convention popularised by documentation platforms, and many sites link to it from the Optional section of their llms.txt.

Next steps