How Do AI Search Engines Decide What to Recommend?
Four stages, each with different consequences for a business. Some of what follows is published by the companies involved. Most of what the industry asserts about it is not. This page keeps the two apart on every claim.
What Actually Happens When Somebody Asks
At plain level four things happen between a question being typed and an answer appearing. Each stage has different implications for a business, which is why they are worth separating rather than treating the whole thing as a black box.
One. The question is interpreted. What was typed becomes what the system will go looking for.
A conversational question gets turned into something the retrieval stage can act on. This is where a long, oddly phrased question can end up matching material that shares none of its wording, which is why writing for exact phrases matters less here than it did.
Two. Sources are retrieved. A small number of candidates are gathered.
Not ten, not a hundred. A handful. Block two is about why that number changes everything. Block three covers where the candidates come from, which differs by product.
Three. An answer is generated. Prose is composed from what was gathered.
The system writes rather than lists. What ends up in the answer depends on what was retrieved and on how usable each candidate was, which is the part nobody outside these companies can describe precisely.
Four. Some sources are named. Not all of them, not always.
Products differ. Some list sources against each statement, some give a short list at the end, some name nothing. A business can be used without being credited, which is a real commercial outcome with no measurement attached to it.
Why the stages matter separately. They fail separately.
A business can be invisible at stage two and nothing else matters. It can be retrieved at stage two and unusable at stage three. It can be used at stage three and uncredited at stage four. Three different problems needing three different responses.
This page is reviewed monthly. Product behaviour changes faster than the underlying idea.
Retrieval Is Not Ranking
Ordering ten results and gathering four sources to write from are different operations with different consequences. Almost every misunderstanding in this subject comes from treating the second as though it were the first.
What ranking does. Sorts a long list.
A results page puts things in order and lets the person choose. Position matters because attention falls away down the page. Being sixth is worth less than being second. Both are worth something.
What retrieval does. Selects a few, then discards the rest entirely.
If four sources are gathered and yours is not among them, you are not sixth. You are absent. The answer is written as though you do not exist. There is no long tail of diminished visibility below the cut.
Why that changes what visibility means. It becomes closer to binary.
The useful question stops being where you place and becomes whether you were included at all. That is a harsher test in one direction and a more forgiving one in another, since a business that gets included is not competing for attention against nine others on the same screen.
What it does to the idea of a position. Removes it.
There is no first place in a set of four sources feeding a paragraph. Even where sources are listed, the order tells you very little about what the answer leaned on. Anybody reporting your position in an assistant is reporting something the system does not produce.
The one thing this makes simpler. The goal.
Instead of improving a position, you are trying to be in the set. That is a clearer objective than it sounds. Per our guide to what you can influence it points the work in a specific direction.
Where The Sources Come From
There is no single answer. The differences matter commercially. A business invisible through one route can be perfectly visible through another, which means one product answering without you tells you less than you might assume.
A search index. The system queries an existing index and works from what comes back.
Where a product does this, ordinary search visibility and this are closely connected. Material that cannot be found conventionally is unlikely to arrive at the retrieval stage at all.
A live fetch. Pages are requested at the moment the question is asked.
This is the route where crawler permissions bite hardest. A site that declines the relevant crawler cannot be fetched, whatever its ordinary search visibility. That is documented behaviour rather than inference, since vendors publish their crawler names and how to permit or refuse them.
The model itself. What was absorbed during training.
Material a model saw in training can inform an answer with nothing fetched at all. This route is the least controllable and the slowest moving. It also cannot be corrected quickly, since a business whose old description was absorbed cannot edit that.
A mixture. Which is what most products actually do.
Several routes combined, weighted in ways nobody outside the companies has published. Treat any confident account of the mixture as inference.
What follows practically. Cover the routes you can influence.
Conventional visibility and crawler access are both actionable. Training data is not. Our four platform guides cover what each company has published, starting with how ChatGPT handles business questions.
What Gets A Source Named
Everything in this block is observation. No company has published how candidates are chosen or weighted at the point an answer is written. What follows is what practitioners see repeatedly, offered as pattern rather than rule.
Relevance to the actual question. The least surprising of the three.
Material that addresses the specific question appears more often than material that addresses the general subject. A page covering a topic broadly seems to lose to a page answering the exact thing asked, which is consistent with how retrieval is described to work.
Stating things clearly and without hedging. The pattern that surprises people.
Passages that make a definite statement appear to be used more readily than passages that circle a subject. This appears to be a property of extractability rather than of quality: a sentence that answers something completely can be lifted, while one that depends on three qualifications cannot.
Being corroborated elsewhere. The one with the most practitioner agreement behind it.
Businesses described consistently across several independent sources appear more often than businesses described in only one place. The reasoning offered for this is that an assembling system has no way to verify a lone claim. We find that plausible. Nobody has published it.
What we will not tell you. How these are weighted against each other.
Whether corroboration outweighs clarity is unknown outside these companies. So is whether relevance outweighs both. Any list presenting weighted factors as fact has invented the weights.
Why observation is still worth acting on. The actions are safe.
Each of the three points at work that helps regardless: answer questions specifically, write clearly, be described consistently. None of it is wasted if the pattern shifts, which is what makes it reasonable to act on incomplete knowledge.
What Is Documented And What Is Inferred
Three categories of claim circulate about this subject. They carry completely different weight and they are presented identically almost everywhere, which is the single biggest problem with published material here.
Documented. Published by the company that operates the product.
Crawler names and how to permit or refuse them. What a product broadly does. General guidance on content. This material is narrow, it is real. It can be cited with a date. Every claim of this kind on our platform pages carries the company and the date in the same sentence.
Inferred. Observed behaviour, reasoned about.
Practitioners ask questions, record what appears and look for patterns. Done carefully this is legitimate and useful. It is not mechanism. A pattern that held in spring may not hold after a model update. Observation also cannot distinguish a cause from a coincidence.
Invented. Presented as fact, sourced to nothing.
Named ranking factors with percentages attached. Scores described as your visibility. Confident accounts of internal weighting. This material exists because there is demand for certainty and no penalty for supplying it falsely.
How to tell them apart. Two questions. They work on anybody.
Who published this, then when. A documented claim survives both. A careful inference is labelled as observation by whoever made it. An invention survives neither, which is why so much published material on this subject avoids attribution entirely.
Apply it to this page. That is the point of the block.
Block three's crawler point is documented. Block four is labelled observation throughout. Where we do not know something, we have said so rather than filling the gap. If a competitor's page on this subject cannot survive the same test, that tells you what their recommendations rest on.
Why The Same Question Gives Different Answers
Ask the same question twice and you can get two different sets of businesses. This is not a fault and it is not something a supplier can tune around. It follows from how these systems work.
Non-determinism. The same input does not guarantee the same output.
Generation involves an element of variability by design. Two identical questions can produce differently worded answers drawing on different sources, with nothing having changed on anybody's website.
Personalisation. The asker affects the answer.
Location, conversation history and in some products a stored memory of earlier exchanges can all shape what comes back. Two people in different towns asking identically are not running the same query.
Model versions. The system underneath changes.
Products are updated on schedules nobody outside publishes in advance. A change on their side can move which businesses get named more than anything a business did to its own site. This is the reason attributing progress here requires care, which our measurement guide deals with.
What this does to position tracking. Removes its meaning rather than its difficulty.
A tracked position assumes a stable ordering that can be sampled. Here there is no ordering and no stability, so a recorded position describes one instance of a variable output. Repeating the check produces a different number without anything having improved or declined.
What to do instead. Sample deliberately and watch the trend.
A consistent set of questions asked on a schedule, with what appears recorded each time, tells you something over months. One check tells you nothing at all.
What This Means For A Business
Four instructions follow from everything above. They are deliberately unexciting, since the mechanism does not support anything clever.
Be findable. Through every route you can influence.
Conventional search visibility, because several products retrieve from an index. Crawler access, because several fetch live. Both are actionable and one of them is documented rather than inferred.
Be clear. Because extractability appears to matter.
Write so a passage answers something completely without depending on the paragraph before it. This is the one piece of writing advice in this subject that we would give even if the whole field evaporated, since it makes pages better for people too.
Be corroborated. Described the same way in more than one place.
Consistent business information across the web, plus genuine references from independent sources. Earned rather than bought, as our backlinks guides have always argued.
Stop expecting a rank. The hardest of the four.
There is no position to hold, so a report showing one is describing something the system does not produce. Judging this work requires different evidence. A supplier who has not told you that has either not understood the mechanism or has decided you would rather not know.
What none of this is. A method for being recommended.
These are conditions that appear to make inclusion more likely. Our pillar guide states the same position, which is that the offer here is better odds rather than an outcome.
If you are not in the four,
you are not sixth.
Retrieval gathers a handful of sources then discards the rest. There is no long tail of diminished visibility below the cut, which makes the goal being included rather than placing well.
What is included every month:
£350 per month, one target area. No setup fee, nothing billed separately.
Twenty guides.
One subject.
The four platforms individually, what you can influence, whether it is worth it yet, what it costs, how to measure it and how to judge a supplier selling it.