The 30-second answer: to get cited by Perplexity, be in its index, be readable as text, and answer the question in the first paragraph
Perplexity answers a question by searching the web, reading the pages it finds, and citing them. Two agents do the work: PerplexityBot builds the search index that candidate pages come from, and Perplexity-User fetches a page at answer time when a question needs it. So the levers are the ones a search engine would reward, with one twist: the page is read by a model that quotes, so the first paragraph has to contain the answer, the number, and the date.
Five steps, in order of effect: allow PerplexityBot and its published IP ranges; make each important page answer one question as text in the HTML; look like a source a reviewer would trust, with named authors and dated pages; keep your facts consistent everywhere; and measure with logs, referrals, and a monthly audit.
How Perplexity finds and cites a page
Perplexity’s help center describes the product as searching the internet for each question and gathering information from authoritative sources such as articles, websites, and journals, then showing the sources it used. It does not publish a ranking method. The bot documentation is more specific about the mechanics:
- PerplexityBot is designed to surface and link websites in search results on Perplexity, and is not used to crawl content for AI foundation models. To appear in results, Perplexity recommends allowing it in robots.txt and permitting requests from its published IP ranges.
- Perplexity-User supports user actions: when a person asks a question, it may visit a page to help provide an accurate answer and include a link to the page in the response. Because a user requested the fetch, it generally ignores robots.txt.
- IP ranges are published as JSON at perplexity.com/perplexitybot.json and perplexity-user.json, each a creation time plus a list of IPv4 prefixes. Robots.txt changes may take up to 24 hours to be reflected.
The consequence for a site owner: there is an index to be in, and there is a fetch at answer time that reads the page as served. Both are plain HTTP requests for HTML. Nothing in the documentation mentions rendering JavaScript, sitemaps, or any file other than robots.txt.
Step 1. Let PerplexityBot in, and verify it is the real one
Allow the bot by name, and do not let a firewall rule for unknown agents undo it:
User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: /
If your CDN or WAF challenges non-browser requests, add the published prefixes to its allow list; Perplexity says WAF changes may take some time to propagate and recommends monitoring logs. Verify requests the same way: a hit claiming to be PerplexityBot from an address outside the JSON is not PerplexityBot. The crawler list has the verification steps for every vendor.
Step 2. Make each important page quotable
The fetch at answer time reads the HTML as served and the model lifts sentences from it. Three habits make a page the one it lifts from:
- One page, one question, answered first. Put the question in the heading and a two-sentence answer with the number and the date in the first paragraph. Explanation follows.
- Text in the HTML. Prices in images, content behind tabs, and sections loaded by script after the page arrives are invisible to a plain fetch.
- Something worth citing. A figure with its source, a quotation, a dated reference. The one controlled study of generated answers found these raised visibility by roughly a third; keyword density did nothing.
Step 3. Look like a source a reviewer would trust
Perplexity’s help center describes source labels: a shield icon that marks a source as Government, Academic, or Trusted, based on a review of the whole website rather than individual pages. The review asks questions such as whether the site corrects its mistakes, whether it says who wrote each piece, and whether it separates news from opinion. The article states that partnerships, payments, and other business arrangements do not affect a label.
You do not need a label to be cited; most cited pages carry none. But the criteria are a free checklist for what a reader, human or model, treats as credible: a named author on each page, a visible publication and update date, an about page that says who runs the site, and a corrections note when you fix something. These cost an afternoon and apply under every assistant, not just this one.
Step 4. Keep facts consistent, and know what llms.txt does here
A real-time engine finds every copy of a fact you have published. If the pricing page says one thing and a 2024 blog post says another, the answer may cite either. Date the current page, add a dated note to old posts, and make structured data match the visible text; the pricing article walks through the sweep.
On llms.txt, be realistic. In the 12-week panel of 83 sites described in this site’s ChatGPT article, Perplexity fetched the file zero times. Write it for the agents and tools that read it, and treat the pages as the Perplexity surface.
Step 5. Measure
- Logs. Count PerplexityBot and Perplexity-User requests per week, verified against the IP JSON. The index crawler should appear on its own schedule; the user fetcher appears when someone asked.
- Referrals. Visits from perplexity.ai in analytics. These are citations that turned into clicks.
- A monthly audit. Ask Perplexity your ten most important questions in a logged-out session. Record whether you are in the source list, your position, and which page was cited. The citations article has the full four-check routine and what normal looks like in month one.
When you are not cited
| Symptom | Likely cause | Fix |
|---|---|---|
| No PerplexityBot in logs at all | robots.txt or a WAF rule blocks it | Allow by name and by IP range; wait up to 24 hours |
| Bot visits, never cited | No page answers the question directly, or the answer is below the fold | One page per question; answer in the first paragraph |
| Cited, but with the wrong number | An old copy of the fact outranks the current page | Date the current page; annotate or redirect old copies |
| Competitor cited for your own brand query | Your title or first paragraph does not name the thing | Put the product and company name in the title and opening |
| Cited on ChatGPT, not Perplexity | Different pipelines; Perplexity favours recent, structured pages | Add dates, update the page, keep tracking separately |
| Perplexity-User hits, page not quoted | Content not in the HTML as served | Server-render the text; drop image-only facts |
The first step is the one that fails silently. A checker fetches your robots.txt and pages the way a bot does and reports what PerplexityBot and the other common agents are allowed to read.
See which bots can reach your site
Run the checker on your URL. It fetches your robots.txt and reports what the common AI user agents are allowed to read, alongside the llms.txt draft it generates.
Run the check →FAQ
How does Perplexity decide which pages to cite?
Perplexity's help center describes the product as searching the internet for each question and gathering information from authoritative sources such as articles, websites, and journals, then showing the sources it used. It does not publish a ranking formula. What is documented is the plumbing: PerplexityBot builds the search index that candidate pages come from, and Perplexity-User fetches a page at answer time when a question needs it. A page that is in the index, readable as text, and answers the question directly is what gets cited.
Do I need to allow PerplexityBot in robots.txt?
Yes. Perplexity's bot documentation says that to ensure your site appears in search results you should allow PerplexityBot in robots.txt, and also permit requests from its published IP ranges. It adds that PerplexityBot is not used to crawl content for AI foundation models, so allowing it is not a training decision. Changes to robots.txt may take up to 24 hours to be reflected.
Does Perplexity read llms.txt?
Not that anyone has observed. In the 12-week panel of 83 sites cited in this site's ChatGPT article, Perplexity fetched llms.txt zero times. Perplexity reads your pages. Write llms.txt for the agents and tools that do use it, and spend the Perplexity effort on the pages themselves.
What are the shield icons next to some sources?
Perplexity's help center describes source labels: a shield that marks a source as Government, Academic, or Trusted, based on a review of the whole website rather than individual pages. The review asks questions such as whether the site corrects its mistakes and whether it says who wrote each piece, and the article states that partnerships, payments, and other business arrangements do not affect a label. A label is not required to be cited; most citations carry none.
Why am I cited on Perplexity but not on ChatGPT, or the other way round?
Different pipelines. Perplexity is closest to a real-time search engine and tends to cite recently published, well-structured pages; ChatGPT's browsing path is more conservative. Track each assistant separately and do not expect parity. The crawler articles on this site describe the user agents each one sends.
How do I know it is working?
Three signals: PerplexityBot and Perplexity-User appearing in your server logs, verified against the published IP ranges; referral traffic from perplexity.ai in analytics; and a monthly audit where you ask Perplexity your ten most important questions in a logged-out session and record whether you appear in the source list and in what position.
Next steps
- → How to check whether AI search has picked up your site (the four checks and the monthly audit)
- → AI crawlers list: which bots to allow in robots.txt (user agents and IP verification for every vendor)
- → How to measure the impact of llms.txt and AI search
- → All Learn articles