Skip to content
All writingHow AI cites

How Perplexity decides what to cite

Perplexity cites more sources per answer than any other major engine and searches the live web every time. What it does not do is prefer fresh content, whatever you have been told.

· 7 min read

Perplexity decides what to cite by running a live web search on every single query, ranking what comes back on relevance, entity clarity and authority, then quoting the passages it can lift without distortion. It cites more sources per answer than any other major engine, a mean of 16.35 in one study and 21.87 in another. The thing most people assume about it, that it rewards recently published pages, does not hold: across 17 million cited URLs, the pages Perplexity cites average 1,166 days old, the oldest of any AI assistant measured.

Perplexity is the engine that shows its work. Every answer arrives with numbered citations, which makes it the easiest engine to study and the least forgiving to fake. It is also the engine most people misread in two directions at once. The instinct is to assume an engine that cites openly must be picky, and that an engine searching live must want the newest thing. The data says it is the least picky of the major engines, and that its citations skew older than ChatGPT's. This post covers what it actually rewards.

How Perplexity turns a question into citations

Perplexity answers by retrieving, not remembering. It runs a real-time web search for every single query, pulls back a candidate set, ranks it, composes an answer from the survivors, and attaches a numbered reference to each claim it makes. Nothing in that chain rewards a page for being famous. It rewards a page for being findable and easy to quote accurately.

Roughly, the sequence is: work out what the question is asking, search the live web for it, rank the candidates, keep what fits, write the answer, cite the pages the answer leans on. That is the general shape of retrieval-augmented search rather than a documented internal architecture, and it is worth holding loosely. What is measurable is the output, and the output is unusually generous. Across 602 prompts and 21,143 citations, Perplexity cited a mean of 16.35 sources per answer, against 12.06 for Google AI Overviews and 6.88 for ChatGPT. A separate analysis puts Perplexity at 21.87 citations per response, the highest of any major AI platform. The samples differ, the direction does not.

That matters for strategy. On ChatGPT you compete for one of about seven slots, so depth and authority win. On Perplexity there are two to three times as many slots, so the binding constraint is rarely whether you are the single best source. It is whether you were retrieved at all.

The freshness myth

It is widely repeated that Perplexity favours recent content, and it is the first thing people try when their citations dry up. The largest dataset available does not support it. Ahrefs studied 16.975 million cited URLs across seven AI platforms and found the pages Perplexity cites average 1,166 days old, roughly three years and two months. That is older than ChatGPT at 958 days, Copilot at 1,056 and Gemini at 1,118.

AI assistants as a group do skew fresher than organic search, 1,064 days against 1,416, so freshness is not irrelevant. But Perplexity is the least fresh of the assistants, not the most, and publishing something new is not the lever that gets you into its citations. If your pages stopped being cited, recency is the wrong first thing to check.

What live retrieval actually buys you is different: a page published today can be cited today, with no waiting for an index to catch up. That is access, not preference. Being new gets you eligible. It does not get you picked.

What Perplexity ranks on

With recency demoted, what is left is the unglamorous list: relevance to the question, entity clarity, and authority. Two of the three are measurable.

Signals Perplexity ranks on: relevance, fact density, entity clarity and authority
Retrieved is not cited. The gap between them is where the work is.

On density, the GEO: Generative Engine Optimization paper from Princeton and IIT Delhi, presented at KDD 2024, found that adding quotations, statistics and explicit sourcing lifted a page's visibility in generative answers by up to 40% across 10,000 queries, while keyword stuffing did not help. On authority, third-party presence does more than your own domain can: domains with active profiles on platforms like G2 or Capterra show roughly three times higher citation probability. And position still matters, because around 44% of AI citations come from the first 30% of a page.

Why Perplexity leans on Reddit and third-party pages

Perplexity leans on third-party pages because a retrieval-based engine treats independent sources as validation rather than marketing. A Reddit thread, a G2 listing or a journalist's roundup reads as what other people concluded, which outranks your own landing page making the same claim. Reddit is the clearest case: it is the most-cited single domain across engines. Perplexity is the notable exception on volume, though. On Semrush's tracking of 217,000 prompts, refreshed October 2025, Reddit appears in just 3.5% of Perplexity answers, against 12.6% of SearchGPT's and 9% of Google AI Mode's. Perplexity leans on third-party pages heavily; it just spreads that weight across more of them than its rivals do.

This also explains the age data. A Reddit thread that has accumulated answers for two years is a better source than one posted this morning, and Perplexity's citation mix is full of them. The off-site tilt is easy to miss if you only watch your own domain: only 38% of the pages cited in Google's AI Overviews also rank in its top 10, down from 76% eight months earlier. This is the blind spot on-site AEO dashboards have, and it is why your AEO tool shows nothing for off-site work. For the full source breakdown, where AI actually gets its answers lays out the data.

How to get cited by Perplexity

Be retrievable and be quotable. Open each section with a direct answer, pack it with specific facts and named entities, and keep sections self-contained so a passage survives being lifted out of context. Because Perplexity cites widely, you do not have to beat every competitor to get in. You have to be found and be easy to quote correctly.

Put the answer high, because 44% of citations come from the first third of a page. Lead with the conclusion, back it with a hard number, name specific tools and platforms rather than "AI tools", and structure with clean headings and FAQ blocks. Then spend the effort you were about to spend on republishing dates on getting into the third-party pages instead, because that is where Perplexity's citations actually live. The same format rules apply inside a Reddit comment or a review, which is why getting cited on Reddit by ChatGPT uses the identical discipline.

How Perplexity differs from ChatGPT and AI Overviews

The engines differ most in breadth. Perplexity cites the most, a mean of 16.35. ChatGPT cites the fewest at 6.88 and extracts about 4.2 times more language from each source it keeps, so depth and authority matter more there. Google AI Overviews sit between them on count and cite the oldest content of the three, averaging 1,432 days, almost exactly matching organic search.

The consequence is that a win on one engine does not transfer. Getting into Perplexity is largely a retrieval problem: exist in the right places, in quotable form. Getting into ChatGPT is a selection problem: be the one source worth quoting at length. How ChatGPT chooses which sources to cite covers the other side of that trade, and how Google AI Overviews and AI Mode pick sources covers the third, where a query fans out and the citation lands on a passage rather than a page. The through-line: measure each engine separately, and measure it off-site.

How to tell whether Perplexity is actually citing you

You tell whether Perplexity is citing you by running your real buyer queries in Perplexity on a schedule and logging every numbered source it returns, then checking how many are yours, how many are third-party pages about you, and how the mix shifts. Perplexity makes this easier than any other engine because it shows its citations openly, but its answers still change, so a one-time check tells you nothing.

Set up a simple loop: freeze a list of 10 to 30 buyer questions, run them in Perplexity and your other target engines on a regular cadence, and record every cited URL so you can separate your own domain from the Reddit threads, reviews and roundups doing the real work. Then track your off-site AI Share of Voice, the share of cited sources that are about you, the way AI Share of Voice, off-site lays out. Because Perplexity cites so many sources per answer, a small change in your off-site footprint moves this number faster here than anywhere else.

Where this leaves you

Perplexity is the most generous citer of the major engines, and the one whose citations skew oldest. It searches live on every query, quotes widely, and most of what it quotes about you sits on domains you do not own. So the winning move is not to publish more often. It is to exist, in quotable form, on the forums, reviews and editorial pages it keeps returning to, and to write passages built to be lifted: answer first, dense with facts, named entities, self-contained. Then measure it per engine, over time, off-site.

Plugged in, or plugged out?

Find out what AI is already quoting about you.

free to start · no credit card