The 30-second answer: most llms.txt mistakes are delivery problems, not writing problems
Of the twelve, four are about the server (the file is not really there, the wrong Content-Type, a bad encoding, bots blocked), four are about Markdown structure (the H1, the blockquote, the link syntax, relative URLs), and four are about content (dead links, a file that is too big, annotations that say nothing, a file nobody maintains). Fetch the file with curl before you read it in a browser: half the list is visible in the first ten lines of the response.
| # | Mistake | Validator check | Level |
|---|---|---|---|
| 1 | The file is an HTML fallback page | content_type_ok | fail |
| 2 | Wrong Content-Type | content_type_ok | warn |
| 3 | Not UTF-8, or a BOM | encoding_utf8 | fail / warn |
| 4 | Bots blocked by robots.txt or a WAF | served_at_root, links_reachable | fail |
| 5 | No H1, several H1s, or H1 not first | has_h1, single_h1, h1_first | fail / warn |
| 6 | No blockquote summary | has_summary | warn |
| 7 | Links not in list format | has_links, link_format_valid | fail / warn |
| 8 | Relative URLs | links_use_http | fail |
| 9 | Dead links | links_reachable | warn / fail |
| 10 | Too big: a sitemap dump | reasonable_size | warn |
| 11 | Annotations that repeat the link text | (human check) | |
| 12 | Nobody owns the file | (human check) |
Delivery mistakes
1. The file is an HTML fallback page
The most common failure. The server answers /llms.txt with status 200, and the body is your single-page app shell or a custom 404 page. A browser shows a normal page; a parser finds a doctype where it expected a heading. The validator checks both the Content-Type header and the first 512 bytes of the body for an HTML document and fails either. Fix: put the file in the static directory your host serves verbatim, or add a route that returns text, then confirm with curl -i https://your-site/llms.txt that the first body line starts with #.
2. Wrong Content-Type
The file is there but served as application/octet-stream or with no type at all, because the host has no mapping for a .txt at that path. Most readers cope, but the header is the first thing a strict client sees. Serve text/markdown; charset=utf-8 or text/plain; charset=utf-8. This site serves its own file as text/markdown.
3. Not UTF-8, or a byte-order mark
A file saved from an editor in a legacy encoding, or with a UTF-8 BOM, produces a first line that is not a heading. The validator fails a body that does not decode as UTF-8 and warns on a BOM, a charset parameter other than utf-8, or replacement characters in the text. Fix: re-save as UTF-8 without BOM and declare charset=utf-8 in the header.
4. Bots blocked by robots.txt or a firewall
The file is perfect and nothing can read it. A robots.txt rule that disallows everything for unknown agents, or a WAF default that challenges non-browser requests, returns a 403 or a challenge page to the crawlers you wrote the file for. The validator reports no response or a non-200 status at the root, and if the file is reachable but every sampled link fails, it says so. Fix: allow /llms.txt and the pages it links to for the agents you want, and test with the crawler user agents rather than a browser.
Structure mistakes
5. No H1, several H1s, or the H1 is not first
The spec has exactly one required element: an H1 with the name of the project or site. Files that start with a comment, a blank blockquote, or a logo line fail the first check a parser makes. Two H1s, usually a site name and then a product name, leave the reader guessing which is the title. Fix: one line, # Site name, as the first non-empty line.
6. No blockquote summary, or a slogan in its place
The spec describes a blockquote with a short summary containing the key information needed to understand the rest of the file. The validator warns when there is none. The variant it cannot catch is a tagline: “Build faster. Ship sooner.” says nothing a model can use. Fix: one or two sentences with the noun (what the thing is), the audience, and the one fact a reader needs before the links.
7. Links that are not list items
Bare URLs on their own lines, a Markdown table, numbered lists, or HTML anchors. The spec defines each entry as a list item with a required hyperlink and optional notes after a colon. The validator fails a file with no valid items and warns when some bullets do not match. The fix is mechanical:
## Docs https://acme.example/docs/quickstart 1. Install guide - https://acme.example/docs/install | Pricing | https://acme.example/pricing | ## Docs - [Quick start](https://acme.example/docs/quickstart): Install and first run in five minutes. - [Install guide](https://acme.example/docs/install): Supported platforms and package managers. - [Pricing](https://acme.example/pricing): Three plans; the free tier covers one project.
8. Relative URLs
[Docs](/docs) works in a browser that knows the base URL and nowhere else. The spec’s examples use absolute URLs, and a file read through a proxy, a cache, or an agent’s saved copy has no base to resolve against. The validator fails any link that is not http or https. Fix: write the scheme and host every time, and generate the file if that is tedious.
Content mistakes
9. Dead links
Pages renamed, docs moved to a new host, a trailing slash that now redirects to a 404. The validator probes a sample of up to ten links with HEAD requests and reports how many answered. Crawlers do not report; they skip the link and move on. Fix: run the check after every site restructure, and prefer linking to stable landing pages over deep URLs that change.
10. Too big: the sitemap dumped into Markdown
A generator that writes one bullet per page produces a file that is complete and useless: hundreds of links, no sections worth the name, and no way for a reader with a few thousand tokens to choose. The validator warns above 100 KiB and stops at 512 KiB. Most well-made files are under 15 KB. Fix: the index lists the ten to fifty pages that answer real questions, grouped under named sections; the full text, if you want to publish it, goes in llms-full.txt.
11. Annotations that repeat the link text
“[Pricing](…): Pricing page.” passes every automated check and tells a model nothing. The notes after the colon are the one place in the file where you can say what a reader will find and who it is for. Fix: one sentence per link with a fact that is not in the title: the plan count, the platforms covered, the question the page answers.
12. Nobody owns the file
The file was written once, by hand, during a launch week. Six months later it links to retired pages, quotes an old price, and describes a product that has been renamed. Nothing in a validator catches a true sentence that stopped being true. Fix: generate the file from the same data as the sitemap where you can, and where you cannot, put it on the checklist for every pricing change, rename, and docs move, with a dated line in the summary so a reader can tell when it was last reviewed.
Two more that the validator flags as cosmetic: a file without a trailing newline, and a file with no ## sections at all. Neither breaks a reader, and both take a minute.
Checking your file
Fetch the file with a plain HTTP client first; then run it through a validator, which reports each of the checks above by name and lists the sections and links it found.
Validate your llms.txt
Paste your site URL and the validator fetches /llms.txt, runs the checks in this article, and lists every section and link it found, with a sample of links probed for reachability.
Open the validator →FAQ
My llms.txt returns 200 but the validator says it is not there. Why?
Because the body is HTML. Single-page apps and some hosts answer every unknown path with the app shell or a custom 404 page and a 200 status, so a fetch succeeds and the content is a web page. The validator checks both the Content-Type header and the first bytes of the body for an HTML document. Add the file to the static assets or a route that serves text, and confirm with curl that the response starts with a # heading.
Does the Content-Type really matter?
The spec says nothing about it, and most readers will parse text served as text/html if the body is Markdown. But a wrong type is usually a symptom: text/html means a fallback page, application/octet-stream means the host does not know the file. Serve text/markdown or text/plain with charset=utf-8 and the symptom disappears with the cause.
Are relative URLs allowed?
The spec's examples use absolute URLs and describe each list item as a hyperlink to where further detail is available. A relative path is ambiguous to a reader that fetched the file through a proxy or saved it elsewhere, and this site's validator fails any link that is not http(s). Write the full URL, including the scheme and host.
How big is too big?
The spec sets no limit. This validator warns above 100 KiB and stops reading at 512 KiB, because a model with a few thousand tokens of budget cannot use more. Most well-made files are under 15 KB. If yours is larger, it is usually a sitemap export or every page of a documentation set; keep the index short and move full text to llms-full.txt.
Does the validator check my annotations?
No. It checks that each bullet matches the link format and that the URL is http(s); the notes after the colon are free text. Whether the note says something a link text does not is a human check, and it is the mistake that survives validation most often.
Do I need to fix warnings, or only failures?
Failures mean a reader will not get a usable file: unreachable, HTML, no H1, no links, not UTF-8. Warnings mean the file works but departs from what the spec recommends or what parsers expect: no blockquote, no sections, bullets in another format, a large size, a missing trailing newline. Fix failures first; most warnings take a minute each.
Next steps
- → How to write llms.txt by hand (the file that avoids all twelve, in 30 minutes)
- → The llms.txt spec, explained line by line (what is required and what is convention)
- → llms.txt examples: 8 real files, annotated
- → All Learn articles