KnownByLLM

Explainer · 9 min read

llms.txt vs schema.org structured data

One describes a page to a crawler. The other describes a site to an assistant. You want both, for different reasons.

Two files are supposed to make your site legible to machines. Structured data, usually schema.org vocabulary in a JSON-LD script tag, has been in pages since 2011. llms.txt, a Markdown file at the site root, arrived in 2024. People ask whether the new one replaces the old one, or whether either matters for AI answers.

This explainer sets the two side by side: what each contains, who reads it and when, what the documentation from Google and the llms.txt spec actually promise, where the two overlap, and how to keep them from contradicting each other.

The 30-second answer: llms.txt vs schema.org is not a choice, because they describe different things to different readers

schema.org structured data is in-page markup. It tells a crawler what the entities on this one page are: this is an Article with this headline and date, this is a Product with this price, this is an Organization with this name and logo. It is read at crawl time, stored in an index, and used to build search features before anyone asks a question.

llms.txt is a site-level index in Markdown. It tells an agent, at the moment it needs to know about your site, what the site is and which pages to fetch for which topic. It can point at pages on other domains and at Markdown versions of your pages, which structured data cannot. It is read on demand, not indexed.

So the question is not which one to adopt. It is whether both say the same thing about your site. The rest of this article is about making sure they do.

What each one is

schema.org structured datallms.txt
Where it livesInside each page, usually a JSON-LD script tag in the headOne Markdown file at /llms.txt (or under a sub-path)
FormatJSON-LD, Microdata, or RDFa using the schema.org vocabularyMarkdown: H1, blockquote, H2 sections of links
ScopeThe entities on that one pageThe whole site, or the part under its path
Who reads itSearch engine crawlers and anything that parses HTMLAgents and tools that fetch the file by its fixed name
WhenAt crawl time, stored in an indexOn demand, when an agent needs information about the site
Can link off-siteOnly as property values (sameAs, url)Yes; the spec lists this as a reason sitemap.xml is not a substitute
Defined bySchema.org, founded by Google, Microsoft, Yahoo and YandexAn open proposal at llmstxt.org, revised in August 2026
If it is missingNo rich results; the page is still indexed from its HTMLAgents fall back to crawling HTML or to search

The row that matters most is the one about timing. Structured data is consumed before the question exists, by a system that visits millions of pages and keeps what it finds. llms.txt is consumed after the question exists, by a model that has a budget of a few thousand tokens and needs to decide which two or three pages to read. The same fact, say your product name, belongs in both, but for opposite reasons.

What structured data does for AI, according to Google

Google’s page on AI features and your website, last updated in December 2025, is unusually direct. There are no additional requirements to appear in AI Overviews or AI Mode. There is no special schema.org structured data to add. You do not need to create new machine-readable files, AI text files, or markup to appear in those features. The same page keeps one standing request: your structured data should match the visible text on the page.

That last sentence is the useful one. Google’s general structured data guidance says the markup must describe the content of the page it is on, and that you should not add structured data about information that is not visible to the user, even if it is accurate. For a model that reads the page after retrieval, consistent markup is a second copy of the facts, in a shape that is easy to parse. Inconsistent markup is a contradiction it has to resolve.

Two further points from the documentation shape what to write:

  • JSON-LD is the recommended format. Google supports JSON-LD, Microdata, and RDFa and recommends the format that is easiest to implement and maintain, which it says is JSON-LD in most cases. One script tag generated from the same data the page renders is hard to get wrong.
  • Some rich results have been withdrawn. Since August 2023, FAQ rich results are shown only for well-known, authoritative government and health websites. FAQPage markup is still valid and still describes the page, but for most sites it no longer earns anything visible in Search.

The practical reading: structured data is for search engines first, and for AI features only in the sense that those features sit on top of the same index. Nothing in Google’s documentation says a model consults your JSON-LD at answer time.

What llms.txt does that structured data cannot

The llms.txt proposal describes itself as a file that provides information to help agents use a website. Three capabilities follow from its design, and none of them is available to in-page markup:

  1. A site-level view. Structured data describes one page at a time. Nothing in schema.org says which ten pages out of your five thousand matter most. llms.txt is exactly that list, with one sentence per link explaining why.
  2. Links beyond the site. The spec explains why sitemap.xml is not a substitute: a sitemap does not include URLs to external sites even though they might help an agent understand the information, and it does not list LLM-readable versions of pages. llms.txt can do both.
  3. Reading at answer time. The spec contrasts its file with robots.txt: robots.txt tells automated tools what access is acceptable, while llms.txt information is used on demand, when an agent needs information about a topic while assisting a user. That is the opposite end of the pipeline from an index.

The spec also mentions the other file directly. It says the llms.txt file can reference structured data markup used on the site, helping LLMs understand how to interpret that information in context. It does not ask you to paste JSON-LD into Markdown; it suggests a link to wherever your markup is explained.

One caution the documentation shares with Google’s: no major search engine has said it reads llms.txt, and Google has said it does not. The readers today are agents, developer tools, and crawlers that fetch the file by its fixed name. The companion article on whether Google uses llms.txt goes through the public statements.

Where they overlap, and how to keep them consistent

Both files end up stating the same handful of facts. Here is a minimal pair for a fictional company. The JSON-LD sits in the head of the home page; the Markdown is the whole llms.txt.

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Acme Analytics",
  "url": "https://acme.example/",
  "logo": "https://acme.example/logo.png",
  "description": "Privacy-first web analytics for small teams, self-hosted or cloud.",
  "sameAs": ["https://github.com/acme-analytics"]
}
</script>
# Acme Analytics

> Privacy-first web analytics for small teams, self-hosted or cloud. No cookies, no personal data stored.

## Docs

- [Quick start](https://acme.example/docs/quickstart): Install the script and see your first pageview in five minutes.
- [Self-hosting guide](https://acme.example/docs/self-host): Docker image, Postgres, and the two environment variables that matter.
- [Pricing](https://acme.example/pricing): Cloud plans by monthly pageviews; self-hosted is free.

## Optional

- [GitHub](https://github.com/acme-analytics): Source, issues, and release notes.
- [Structured data on this site](https://acme.example/docs/structured-data): Which schema.org types each page type carries.

Three things line up on purpose. The Organization name is the H1. The description is the blockquote, word for word. The sameAs link is the GitHub entry under Optional. When the company renames a plan, both files change in the same commit, and the easiest way to guarantee that is to generate both from one source. The build-time generation article shows the pattern for the Markdown side; the JSON-LD side is the same idea with a different serializer.

The pairs worth checking on your own site:

  • Organization name and description against the llms.txt H1 and blockquote.
  • Article headline and datePublished against the link text and annotation for that page in llms.txt. If the file says a guide is current and the markup says 2023, one of them is wrong.
  • Product name and price against whatever the pricing page link says. Do not put prices in llms.txt annotations unless you regenerate the file when they change.
  • FAQPage questions against the FAQ section of the page they are on. This site renders its article FAQs and the FAQPage JSON-LD from one array for exactly this reason.

What to do this week

  1. Open your home page source and look for application/ld+json. If there is none, add an Organization block with name, url, logo, and description. If there is one, read it and confirm every value is visible on the page.
  2. On your main content type, add Article or Product markup with the properties Google lists as recommended (for Article: author, datePublished, dateModified, headline, image). Skip FAQPage unless the page shows an FAQ.
  3. Write llms.txt by hand if you have under fifty pages, or generate it from the same route list as your sitemap if you have more. Copy the Organization description into the blockquote.
  4. Add one line under Optional pointing at the page that explains your structured data, if you have such a page. If you do not, skip it; the spec allows the reference, it does not require it.
  5. Put both checks in your deploy pipeline: Google’s Rich Results Test or the schema.org validator for the markup, and an llms.txt validator for the file.

Checking both

Structured data has had validators for years; the schema.org validator and Google’s Rich Results Test both parse a URL and list the entities found. For the Markdown side, fetch the file and confirm it has an H1, a blockquote, and named sections whose links resolve.

Validate your llms.txt

Paste your site URL and the validator fetches /llms.txt, checks the structure against the spec, and lists every section and link it found, so you can compare them with what your markup says.

Open the validator →

FAQ

Does llms.txt replace schema.org structured data?

No. They work at different layers. Structured data is in-page markup that describes the entities on one page (an article, a product, an organization) to crawlers that index it. llms.txt is a site-level Markdown index read by an agent when it needs information about your site. One cannot express what the other does, so a site that wants both search features and good answers from assistants keeps both.

Does Google use llms.txt or schema.org for AI Overviews?

Google's guidance for AI features in Search, last updated December 2025, says there are no additional requirements to appear in AI Overviews or AI Mode, no special schema.org types to add, and no need to create new machine-readable files or AI text files. Structured data is still used for rich results and for understanding pages, and Google asks that it match the visible text. Google has said separately that it does not use llms.txt.

Should I put FAQPage markup on every page?

Only where the page actually shows those questions and answers. Google's documentation says FAQ rich results are now shown only for well-known, authoritative government and health websites, so for most sites FAQPage no longer earns a rich result. It still describes the page accurately, and this site keeps it on its long-form articles for that reason. Never add it to a page that does not display the FAQ.

Can llms.txt link to my structured data?

The spec says the file can reference structured data markup used on the site, to help a model interpret that information in context. In practice that means a line in a named section pointing at the page or document that explains your markup, not pasting JSON-LD into the Markdown file. The llms.txt file stays a readable index.

Which format should structured data use?

JSON-LD. Google supports JSON-LD, Microdata, and RDFa and recommends the format that is easiest to implement and maintain, which it says is JSON-LD in most cases. A single script tag in the head, generated from the same data the page renders, is the least error-prone setup and the one this site uses.

If I can only do one this week, which one?

Fix whichever is wrong. A site with no structured data at all should add Organization and Article (or Product) markup first, because every search engine reads it today. A site that already has clean markup but no llms.txt should write the file, because it takes an hour and it is the only one of the two that an assistant can read as a whole.

Next steps