KnownByLLM

Guide · 9 min read

How to get your site cited by Perplexity

Perplexity searches at question time and cites what it reads. Five steps to be what it reads, and a table for when you are not.

Of the major assistants, Perplexity is the one that behaves most like a search engine: every answer comes with a numbered list of sources, and the sources are pages it fetched for that question. That makes it the easiest assistant to reason about, and the one where a small site can see results within days.

This guide covers what Perplexity documents about how it finds and reads pages, the five steps that follow from it, what the source labels are and are not, and a diagnostic table for the common reasons a site is missing from the citation list.

The 30-second answer: to get cited by Perplexity, be in its index, be readable as text, and answer the question in the first paragraph

Perplexity answers a question by searching the web, reading the pages it finds, and citing them. Two agents do the work: PerplexityBot builds the search index that candidate pages come from, and Perplexity-User fetches a page at answer time when a question needs it. So the levers are the ones a search engine would reward, with one twist: the page is read by a model that quotes, so the first paragraph has to contain the answer, the number, and the date.

Five steps, in order of effect: allow PerplexityBot and its published IP ranges; make each important page answer one question as text in the HTML; look like a source a reviewer would trust, with named authors and dated pages; keep your facts consistent everywhere; and measure with logs, referrals, and a monthly audit.

How Perplexity finds and cites a page

Perplexity’s help center describes the product as searching the internet for each question and gathering information from authoritative sources such as articles, websites, and journals, then showing the sources it used. It does not publish a ranking method. The bot documentation is more specific about the mechanics:

  • PerplexityBot is designed to surface and link websites in search results on Perplexity, and is not used to crawl content for AI foundation models. To appear in results, Perplexity recommends allowing it in robots.txt and permitting requests from its published IP ranges.
  • Perplexity-User supports user actions: when a person asks a question, it may visit a page to help provide an accurate answer and include a link to the page in the response. Because a user requested the fetch, it generally ignores robots.txt.
  • IP ranges are published as JSON at perplexity.com/perplexitybot.json and perplexity-user.json, each a creation time plus a list of IPv4 prefixes. Robots.txt changes may take up to 24 hours to be reflected.

The consequence for a site owner: there is an index to be in, and there is a fetch at answer time that reads the page as served. Both are plain HTTP requests for HTML. Nothing in the documentation mentions rendering JavaScript, sitemaps, or any file other than robots.txt.

Step 1. Let PerplexityBot in, and verify it is the real one

Allow the bot by name, and do not let a firewall rule for unknown agents undo it:

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

If your CDN or WAF challenges non-browser requests, add the published prefixes to its allow list; Perplexity says WAF changes may take some time to propagate and recommends monitoring logs. Verify requests the same way: a hit claiming to be PerplexityBot from an address outside the JSON is not PerplexityBot. The crawler list has the verification steps for every vendor.

Step 2. Make each important page quotable

The fetch at answer time reads the HTML as served and the model lifts sentences from it. Three habits make a page the one it lifts from:

  • One page, one question, answered first. Put the question in the heading and a two-sentence answer with the number and the date in the first paragraph. Explanation follows.
  • Text in the HTML. Prices in images, content behind tabs, and sections loaded by script after the page arrives are invisible to a plain fetch.
  • Something worth citing. A figure with its source, a quotation, a dated reference. The one controlled study of generated answers found these raised visibility by roughly a third; keyword density did nothing.

Step 3. Look like a source a reviewer would trust

Perplexity’s help center describes source labels: a shield icon that marks a source as Government, Academic, or Trusted, based on a review of the whole website rather than individual pages. The review asks questions such as whether the site corrects its mistakes, whether it says who wrote each piece, and whether it separates news from opinion. The article states that partnerships, payments, and other business arrangements do not affect a label.

You do not need a label to be cited; most cited pages carry none. But the criteria are a free checklist for what a reader, human or model, treats as credible: a named author on each page, a visible publication and update date, an about page that says who runs the site, and a corrections note when you fix something. These cost an afternoon and apply under every assistant, not just this one.

Step 4. Keep facts consistent, and know what llms.txt does here

A real-time engine finds every copy of a fact you have published. If the pricing page says one thing and a 2024 blog post says another, the answer may cite either. Date the current page, add a dated note to old posts, and make structured data match the visible text; the pricing article walks through the sweep.

On llms.txt, be realistic. In the 12-week panel of 83 sites described in this site’s ChatGPT article, Perplexity fetched the file zero times. Write it for the agents and tools that read it, and treat the pages as the Perplexity surface.

Step 5. Measure

  1. Logs. Count PerplexityBot and Perplexity-User requests per week, verified against the IP JSON. The index crawler should appear on its own schedule; the user fetcher appears when someone asked.
  2. Referrals. Visits from perplexity.ai in analytics. These are citations that turned into clicks.
  3. A monthly audit. Ask Perplexity your ten most important questions in a logged-out session. Record whether you are in the source list, your position, and which page was cited. The citations article has the full four-check routine and what normal looks like in month one.

When you are not cited

SymptomLikely causeFix
No PerplexityBot in logs at allrobots.txt or a WAF rule blocks itAllow by name and by IP range; wait up to 24 hours
Bot visits, never citedNo page answers the question directly, or the answer is below the foldOne page per question; answer in the first paragraph
Cited, but with the wrong numberAn old copy of the fact outranks the current pageDate the current page; annotate or redirect old copies
Competitor cited for your own brand queryYour title or first paragraph does not name the thingPut the product and company name in the title and opening
Cited on ChatGPT, not PerplexityDifferent pipelines; Perplexity favours recent, structured pagesAdd dates, update the page, keep tracking separately
Perplexity-User hits, page not quotedContent not in the HTML as servedServer-render the text; drop image-only facts

The first step is the one that fails silently. A checker fetches your robots.txt and pages the way a bot does and reports what PerplexityBot and the other common agents are allowed to read.

See which bots can reach your site

Run the checker on your URL. It fetches your robots.txt and reports what the common AI user agents are allowed to read, alongside the llms.txt draft it generates.

Run the check →

FAQ

How does Perplexity decide which pages to cite?

Perplexity's help center describes the product as searching the internet for each question and gathering information from authoritative sources such as articles, websites, and journals, then showing the sources it used. It does not publish a ranking formula. What is documented is the plumbing: PerplexityBot builds the search index that candidate pages come from, and Perplexity-User fetches a page at answer time when a question needs it. A page that is in the index, readable as text, and answers the question directly is what gets cited.

Do I need to allow PerplexityBot in robots.txt?

Yes. Perplexity's bot documentation says that to ensure your site appears in search results you should allow PerplexityBot in robots.txt, and also permit requests from its published IP ranges. It adds that PerplexityBot is not used to crawl content for AI foundation models, so allowing it is not a training decision. Changes to robots.txt may take up to 24 hours to be reflected.

Does Perplexity read llms.txt?

Not that anyone has observed. In the 12-week panel of 83 sites cited in this site's ChatGPT article, Perplexity fetched llms.txt zero times. Perplexity reads your pages. Write llms.txt for the agents and tools that do use it, and spend the Perplexity effort on the pages themselves.

What are the shield icons next to some sources?

Perplexity's help center describes source labels: a shield that marks a source as Government, Academic, or Trusted, based on a review of the whole website rather than individual pages. The review asks questions such as whether the site corrects its mistakes and whether it says who wrote each piece, and the article states that partnerships, payments, and other business arrangements do not affect a label. A label is not required to be cited; most citations carry none.

Why am I cited on Perplexity but not on ChatGPT, or the other way round?

Different pipelines. Perplexity is closest to a real-time search engine and tends to cite recently published, well-structured pages; ChatGPT's browsing path is more conservative. Track each assistant separately and do not expect parity. The crawler articles on this site describe the user agents each one sends.

How do I know it is working?

Three signals: PerplexityBot and Perplexity-User appearing in your server logs, verified against the published IP ranges; referral traffic from perplexity.ai in analytics; and a monthly audit where you ask Perplexity your ten most important questions in a logged-out session and record whether you appear in the source list and in what position.

Next steps