Generative Engine Optimisation · Guide

How Do AI Search Engines Decide What to Recommend?

Four stages, each with different consequences for a business. Some of what follows is published by the companies involved. Most of what the industry asserts about it is not. This page keeps the two apart on every claim.

Updated: July 2026
Written by: Andrew Odgers, Managing Director
Reading time: 12 minutes
Four stages

What Actually Happens When Somebody Asks

At plain level four things happen between a question being typed and an answer appearing. Each stage has different implications for a business, which is why they are worth separating rather than treating the whole thing as a black box.

One. The question is interpreted. What was typed becomes what the system will go looking for.

A conversational question gets turned into something the retrieval stage can act on. This is where a long, oddly phrased question can end up matching material that shares none of its wording, which is why writing for exact phrases matters less here than it did.

Two. Sources are retrieved. A small number of candidates are gathered.

Not ten, not a hundred. A handful. Block two is about why that number changes everything. Block three covers where the candidates come from, which differs by product.

Three. An answer is generated. Prose is composed from what was gathered.

The system writes rather than lists. What ends up in the answer depends on what was retrieved and on how usable each candidate was, which is the part nobody outside these companies can describe precisely.

Four. Some sources are named. Not all of them, not always.

Products differ. Some list sources against each statement, some give a short list at the end, some name nothing. A business can be used without being credited, which is a real commercial outcome with no measurement attached to it.

Why the stages matter separately. They fail separately.

A business can be invisible at stage two and nothing else matters. It can be retrieved at stage two and unusable at stage three. It can be used at stage three and uncredited at stage four. Three different problems needing three different responses.

This page is reviewed monthly. Product behaviour changes faster than the underlying idea.

The most useful distinction here

Retrieval Is Not Ranking

Ordering ten results and gathering four sources to write from are different operations with different consequences. Almost every misunderstanding in this subject comes from treating the second as though it were the first.

What ranking does. Sorts a long list.

A results page puts things in order and lets the person choose. Position matters because attention falls away down the page. Being sixth is worth less than being second. Both are worth something.

What retrieval does. Selects a few, then discards the rest entirely.

If four sources are gathered and yours is not among them, you are not sixth. You are absent. The answer is written as though you do not exist. There is no long tail of diminished visibility below the cut.

Why that changes what visibility means. It becomes closer to binary.

The useful question stops being where you place and becomes whether you were included at all. That is a harsher test in one direction and a more forgiving one in another, since a business that gets included is not competing for attention against nine others on the same screen.

What it does to the idea of a position. Removes it.

There is no first place in a set of four sources feeding a paragraph. Even where sources are listed, the order tells you very little about what the answer leaned on. Anybody reporting your position in an assistant is reporting something the system does not produce.

The one thing this makes simpler. The goal.

Instead of improving a position, you are trying to be in the set. That is a clearer objective than it sounds. Per our guide to what you can influence it points the work in a specific direction.

It differs by product

Where The Sources Come From

There is no single answer. The differences matter commercially. A business invisible through one route can be perfectly visible through another, which means one product answering without you tells you less than you might assume.

A search index. The system queries an existing index and works from what comes back.

Where a product does this, ordinary search visibility and this are closely connected. Material that cannot be found conventionally is unlikely to arrive at the retrieval stage at all.

A live fetch. Pages are requested at the moment the question is asked.

This is the route where crawler permissions bite hardest. A site that declines the relevant crawler cannot be fetched, whatever its ordinary search visibility. That is documented behaviour rather than inference, since vendors publish their crawler names and how to permit or refuse them.

The model itself. What was absorbed during training.

Material a model saw in training can inform an answer with nothing fetched at all. This route is the least controllable and the slowest moving. It also cannot be corrected quickly, since a business whose old description was absorbed cannot edit that.

A mixture. Which is what most products actually do.

Several routes combined, weighted in ways nobody outside the companies has published. Treat any confident account of the mixture as inference.

What follows practically. Cover the routes you can influence.

Conventional visibility and crawler access are both actionable. Training data is not. Our four platform guides cover what each company has published, starting with how ChatGPT handles business questions.

Observed throughout, confirmed nowhere

What Gets A Source Named

Everything in this block is observation. No company has published how candidates are chosen or weighted at the point an answer is written. What follows is what practitioners see repeatedly, offered as pattern rather than rule.

Relevance to the actual question. The least surprising of the three.

Material that addresses the specific question appears more often than material that addresses the general subject. A page covering a topic broadly seems to lose to a page answering the exact thing asked, which is consistent with how retrieval is described to work.

Stating things clearly and without hedging. The pattern that surprises people.

Passages that make a definite statement appear to be used more readily than passages that circle a subject. This appears to be a property of extractability rather than of quality: a sentence that answers something completely can be lifted, while one that depends on three qualifications cannot.

Being corroborated elsewhere. The one with the most practitioner agreement behind it.

Businesses described consistently across several independent sources appear more often than businesses described in only one place. The reasoning offered for this is that an assembling system has no way to verify a lone claim. We find that plausible. Nobody has published it.

What we will not tell you. How these are weighted against each other.

Whether corroboration outweighs clarity is unknown outside these companies. So is whether relevance outweighs both. Any list presenting weighted factors as fact has invented the weights.

Why observation is still worth acting on. The actions are safe.

Each of the three points at work that helps regardless: answer questions specifically, write clearly, be described consistently. None of it is wasted if the pattern shifts, which is what makes it reasonable to act on incomplete knowledge.

The line, drawn explicitly

What Is Documented And What Is Inferred

Three categories of claim circulate about this subject. They carry completely different weight and they are presented identically almost everywhere, which is the single biggest problem with published material here.

Documented. Published by the company that operates the product.

Crawler names and how to permit or refuse them. What a product broadly does. General guidance on content. This material is narrow, it is real. It can be cited with a date. Every claim of this kind on our platform pages carries the company and the date in the same sentence.

Inferred. Observed behaviour, reasoned about.

Practitioners ask questions, record what appears and look for patterns. Done carefully this is legitimate and useful. It is not mechanism. A pattern that held in spring may not hold after a model update. Observation also cannot distinguish a cause from a coincidence.

Invented. Presented as fact, sourced to nothing.

Named ranking factors with percentages attached. Scores described as your visibility. Confident accounts of internal weighting. This material exists because there is demand for certainty and no penalty for supplying it falsely.

How to tell them apart. Two questions. They work on anybody.

Who published this, then when. A documented claim survives both. A careful inference is labelled as observation by whoever made it. An invention survives neither, which is why so much published material on this subject avoids attribution entirely.

Apply it to this page. That is the point of the block.

Block three's crawler point is documented. Block four is labelled observation throughout. Where we do not know something, we have said so rather than filling the gap. If a competitor's page on this subject cannot survive the same test, that tells you what their recommendations rest on.

Three reasons, one consequence

Why The Same Question Gives Different Answers

Ask the same question twice and you can get two different sets of businesses. This is not a fault and it is not something a supplier can tune around. It follows from how these systems work.

Non-determinism. The same input does not guarantee the same output.

Generation involves an element of variability by design. Two identical questions can produce differently worded answers drawing on different sources, with nothing having changed on anybody's website.

Personalisation. The asker affects the answer.

Location, conversation history and in some products a stored memory of earlier exchanges can all shape what comes back. Two people in different towns asking identically are not running the same query.

Model versions. The system underneath changes.

Products are updated on schedules nobody outside publishes in advance. A change on their side can move which businesses get named more than anything a business did to its own site. This is the reason attributing progress here requires care, which our measurement guide deals with.

What this does to position tracking. Removes its meaning rather than its difficulty.

A tracked position assumes a stable ordering that can be sampled. Here there is no ordering and no stability, so a recorded position describes one instance of a variable output. Repeating the check produces a different number without anything having improved or declined.

What to do instead. Sample deliberately and watch the trend.

A consistent set of questions asked on a schedule, with what appears recorded each time, tells you something over months. One check tells you nothing at all.

The practical translation

What This Means For A Business

Four instructions follow from everything above. They are deliberately unexciting, since the mechanism does not support anything clever.

Be findable. Through every route you can influence.

Conventional search visibility, because several products retrieve from an index. Crawler access, because several fetch live. Both are actionable and one of them is documented rather than inferred.

Be clear. Because extractability appears to matter.

Write so a passage answers something completely without depending on the paragraph before it. This is the one piece of writing advice in this subject that we would give even if the whole field evaporated, since it makes pages better for people too.

Be corroborated. Described the same way in more than one place.

Consistent business information across the web, plus genuine references from independent sources. Earned rather than bought, as our backlinks guides have always argued.

Stop expecting a rank. The hardest of the four.

There is no position to hold, so a report showing one is describing something the system does not produce. Judging this work requires different evidence. A supplier who has not told you that has either not understood the mechanism or has decided you would rather not know.

What none of this is. A method for being recommended.

These are conditions that appear to make inclusion more likely. Our pillar guide states the same position, which is that the offer here is better odds rather than an outcome.

Generative engine optimisation

If you are not in the four,
you are not sixth.

Retrieval gathers a handful of sources then discards the rest. There is no long tail of diminished visibility below the cut, which makes the goal being included rather than placing well.

What is included every month:

Technical accessibility Crawler access checks Entity consistency Content restructuring Earned references Website management Quarterly technical audits Monthly reporting

£350 per month, one target area. No setup fee, nothing billed separately.

The full guide series

Twenty guides.
One subject.

The four platforms individually, what you can influence, whether it is worth it yet, what it costs, how to measure it and how to judge a supplier selling it.

Questions people ask

How These Systems Decide

How does an AI assistant decide which businesses to mention?
Four stages. The question is interpreted, a small number of sources are retrieved, an answer is written from them, then some sources are named. No company has published how candidates are chosen or weighted at the point the answer is written, so anybody describing that precisely is inferring. What is observed is that relevant, clearly stated and corroborated material appears more often.
Is there a position or ranking we can hold?
No. Retrieval gathers a handful of sources then discards the rest, so if you are not among them you are absent rather than placed low. There is no first place in a set of four sources feeding a paragraph. Anybody reporting your position in an assistant is reporting something the system does not produce.
Why do two people asking the same question get different businesses?
Three reasons. Generation involves variability by design, so the same input does not guarantee the same output. Location, conversation history and stored memory can shape the answer. And the underlying model changes on schedules nobody publishes in advance. A change on their side can move which businesses appear more than anything you did to your site.
How do I tell a real claim from an invented one?
Ask who published it and when. A documented claim survives both questions, since companies do publish crawler names, broadly what their products do and general content guidance. A careful inference is labelled as observation by whoever made it. An invention survives neither, which is why so much published material on this subject avoids attribution entirely.
Does ordinary search visibility still matter for this?
For the products that retrieve from a search index, closely. Material that cannot be found conventionally is unlikely to reach the retrieval stage at all. Other products fetch pages live, where crawler permissions matter most. Some draw on what the underlying model absorbed in training, which is the route you cannot influence. Most products appear to combine them.
So what should a business actually do?
Be findable through the routes you can influence, meaning conventional visibility and crawler access. Write so a passage answers something completely without depending on the paragraph before it. Be described consistently across independent sources. And stop expecting a rank, since there is not one. These are conditions that improve the odds rather than a method for being recommended.