Generative Engine Optimisation · Guide

How Does Perplexity AI Rank and Recommend Businesses?

The one product in this cluster where a business can actually watch this happening. It names its sources consistently, which makes it the closest thing available to a window onto how source selection behaves.

Vendor documentation checked: 30 July 2026
Written by: Andrew Odgers, Managing Director
Reading time: 10 minutes
Sources on the page

What The Product Does Differently

Perplexity presents its answers with the sources listed alongside them. Other products in this cluster do that inconsistently or not at all. That single design choice is why this page exists as a separate guide rather than as a paragraph elsewhere.

Why visible sources change the situation for a business. They convert guesswork into checking.

Everywhere else in this subject you are reasoning about an invisible process. Here you can ask a question and see which businesses were drawn on. That does not tell you why they were chosen. It tells you that they were, which is more than any other product in this cluster gives you.

What the vendor says the product is for. Documented rather than inferred.

Perplexity's crawler documentation, checked on 30 July 2026, describes its main crawler as designed to surface and link websites in Perplexity search results. The same documentation states that it is not used to crawl content for training AI foundation models.

Why that second half is worth noticing. It separates two decisions most businesses conflate.

A business uneasy about its content being used to train models is being told, by the vendor, that this particular crawler is not doing that. Whether to permit it is therefore a narrower question than the general one about AI crawlers.

This page is reviewed monthly. Along with the other three platform guides.

The genuinely useful block

You Can Actually Observe This One

Because sources are shown, a business can run its own observation rather than buying somebody's interpretation of one. This is the only place in the subject where that is true, which makes it worth doing properly.

What to ask. The questions your customers actually ask.

Not your business name, which tells you very little. The questions somebody uses when they do not yet know who to call: the service plus the town, the comparison between two options, the question about cost or timing that precedes an enquiry.

What to record. Which businesses are named. Which sources are cited.

Those are two different things. Sometimes a business is named because a directory listing it was cited rather than because its own site was. That distinction matters, since one is visibility you own and the other is visibility you are borrowing from a third party.

What it tells you. Three things.

Whether anybody in your category is being drawn on at all, which in some categories is the surprise. Which kinds of source the system reaches for, meaning your own site against directories against publishers. And whether that changes over months.

What it does not tell you. Rather more.

It does not tell you why. It does not tell you what would change the outcome. And a single check tells you nothing at all, because answers vary between askings. One observation is an anecdote. A consistent set of questions asked monthly is evidence.

Why we do this rather than buy a score. It is the real thing.

A supplier reporting a visibility figure has done something like this and then converted it into a number with a method you cannot inspect. Our measurement guide explains why we would rather show you the sampling than a score derived from it.

Documented, with dates

Where It Gets Its Sources

Two routes, both described by the vendor. They behave differently enough that the distinction matters when something is not working.

The index. Built by automatic crawling.

Perplexity's documentation, checked on 30 July 2026, describes a crawler that gathers and indexes information to surface and link websites in search results. It recommends that site owners allow it, plus permit requests from the address ranges Perplexity publishes.

The live fetch. Triggered by somebody asking.

The same documentation, checked on 30 July 2026, describes a separate user agent that supports user actions, noting that when somebody asks a question the product may visit a page to help answer accurately and include a link to it in the response.

The candid part. Unusual, plus worth knowing.

Perplexity's documentation, checked on 30 July 2026, states that because a user requested the fetch, that second agent generally ignores instructions in a site's robots file. Very few vendors say anything so direct about the limits of their own controls.

What follows from that. The controls are not symmetrical.

A site can decline the automatic indexing crawler and still be fetched when a person asks about it directly. Anybody telling you that a single instruction removes you from this product entirely has not read the vendor's own documentation.

The settings are separate. Per the same documentation, checked on 30 July 2026, each works independently. Changes can take up to twenty four hours to be reflected.

Pattern, not rule

What Appears To Get A Source Used

Everything here is observation, including our own. Perplexity has not published how it selects or weighs candidate sources when composing an answer, so nobody outside the company can state the mechanism.

Pages that answer one question completely. The most consistent pattern.

A page addressing a single question thoroughly appears to be reached for more often than a broad page mentioning that question among nine others. This is consistent across every product we watch, which is part of why we treat it as the safest thing to act on.

Direct statements rather than hedged ones. Extractability rather than quality.

A sentence that answers something on its own can be lifted. One that depends on three qualifications cannot. That appears to matter more than how well the page is written overall.

Being referenced elsewhere. Corroboration again.

Businesses described consistently by several independent sources appear more often than businesses described only by themselves. The same pattern shows up on every platform in this cluster.

What the observability adds here. You can check these yourself.

On this product you are not taking our word for the three patterns above. You can ask questions in your own category and see whether the sources being cited share those properties. We would rather you did that than trusted us.

What nobody can tell you. The weighting.

Which of the three matters most is unpublished. Any list assigning them percentages has invented the numbers.

Concrete and checkable

Crawler Access

Same treatment as the other platform guides, because this is the part that is documented rather than inferred and the part most often broken by accident.

What the vendor recommends. Explicitly stated.

Perplexity's documentation, checked on 30 July 2026, recommends allowing its indexing crawler in a site's robots file and permitting requests from the address ranges it publishes, so that a site appears in search results.

The address ranges are the part people miss. Permission in one place is not permission everywhere.

A site can allow the crawler in its robots file while a firewall in front of the site refuses the requests anyway. The instruction and the actual traffic are two different things. Only one of them can be confirmed by looking at the site’s own configuration.

The vendor addresses this directly. Which tells you how common it is.

Perplexity's documentation, checked on 30 July 2026, includes guidance for site owners whose web application firewall may be blocking its crawlers, covering how to permit them by matching both the user agent and the published addresses. A vendor writing that section has met the problem repeatedly.

Why this is where we start. It is the failure nobody sees.

Nothing on the site looks wrong. The business simply never appears. No amount of content work will change that while the requests are being refused. We confirm the traffic in server logs rather than assuming the instruction is enough.

And we re-check. After any hosting, firewall or security change.

The transferable value

What This Tells You About The Others

Watching a system that shows its sources is the closest available substitute for watching the ones that do not. That is worth something. It is worth less than people assume.

What transfers. The properties of usable material.

If pages answering one question completely are being cited here, that is evidence about what makes material extractable in general rather than about this product specifically. The same is true of direct statements and of corroboration. Those properties are not product features.

What does not transfer. Anything about selection.

Which sources a different product would have chosen, in what order and for what reason, cannot be inferred from watching this one. Different index, different retrieval, different tuning, none of it published.

The specific error to avoid. Treating one observable system as a proxy for all of them.

A supplier reporting on your visibility across AI search, having actually only watched the product that shows its sources, is generalising from a sample of one. Ask which products a report is based on. If the answer is one, the report describes one.

How we use it. As the calibration, not the measurement.

We watch this product because it is watchable, we treat what we learn as evidence about material rather than about mechanisms. We say which product an observation came from. Our guide to how these systems assemble an answer covers why that discipline matters.

Stated plainly

What Is Not Known

Showing sources is not the same as explaining them. The visibility this product offers is real and it stops at the outcome.

How candidates are selected. Not published.

Perplexity documents what its crawlers are for. It has not published how, from everything available, a handful of sources are chosen for a given answer.

How they are weighted. Not published.

Where several sources are cited, nothing indicates which carried the answer and which was supporting. The list is not a ranking. Treating the first as the winner is reading something into it that is not there.

Why your business was omitted. Not available.

This is the question every client asks and there is no diagnostic. You can see that you were absent. Nothing tells you whether it was accessibility, the material itself or simply variability between askings.

Whether what you observe today holds. Not guaranteed.

Patterns visible now may not survive a model or index change. Those arrive without notice. That is the argument for sampling over months rather than concluding from a week.

What we do with all that. Act on what is safe.

Everything the observation supports doing is work that helps regardless: material that answers questions, accessible to crawlers, corroborated elsewhere. None of it is wasted if the patterns move, which is the only responsible way to act on incomplete knowledge.

Generative engine optimisation

One system you can
actually watch.

Because sources are shown, you can ask questions in your own category and see who is drawn on. We would rather show you that sampling than convert it into a score with a method you cannot inspect.

What is included every month:

Crawler access checks Firewall verification Question sampling Entity consistency Content restructuring Earned references Quarterly technical audits Monthly reporting

£350 per month, one target area. No setup fee, nothing billed separately.

The full guide series

Twenty guides.
One subject.

The mechanism, all four platforms, what you can influence, whether it is worth it yet, what it costs, how to measure it and how to judge a supplier selling it.

Questions people ask

Perplexity and Business Recommendations

Why is Perplexity worth paying attention to specifically?
Because it names its sources consistently, which makes it the only product in this cluster where a business can watch source selection happening rather than reasoning about an invisible process. You can ask questions in your own category and see which businesses were drawn on. It does not tell you why they were chosen. It tells you that they were.
Can we check whether we are being used?
Yes. It needs doing properly. Ask the questions your customers ask rather than your own business name, record which businesses are named and which sources are cited, then repeat on a schedule. Those last two are different: sometimes a business appears because a directory listing it was cited rather than its own site. One check tells you nothing, since answers vary between askings.
Does Perplexity use our content to train AI models?
Perplexity's crawler documentation, checked on 30 July 2026, states that its indexing crawler is designed to surface and link websites in search results and is not used to crawl content for training AI foundation models. It says the same of its user-triggered agent. So whether to permit these crawlers is a narrower question than the general concern about AI crawlers.
If we block the crawler, are we removed entirely?
Not necessarily. Perplexity's documentation, checked on 30 July 2026, describes two separate agents whose settings work independently. It states that because a user requested the fetch, the user-triggered agent generally ignores instructions in a site’s robots file. So a site can decline automatic indexing and still be fetched when somebody asks about it directly.
Our site allows the crawler but we never appear. Why?
A common cause is that the instruction and the actual traffic are different things. A firewall in front of the site can refuse the requests even though the robots file permits them. Perplexity's documentation, checked on 30 July 2026, includes guidance for site owners whose web application firewall may be blocking its crawlers, which tells you how often this happens. Confirming the traffic in server logs is the only way to know.
Can what we learn here be applied to ChatGPT or Gemini?
Partly. What transfers is evidence about which material is usable: pages answering one question completely, direct statements, being corroborated elsewhere. Those are properties of content rather than product features. What does not transfer is anything about selection, since the index, retrieval and tuning differ and none of it is published. A supplier reporting on AI visibility generally while only watching this product is generalising from one.