Generative Engine Optimisation · Guide

How to Choose a GEO Agency in the UK

No standards, no certification and no metric a buyer can verify, which makes this an unusually easy field to sell nonsense in. This page gives you the questions that would expose a weak supplier, including us.

Updated: July 2026
Written by: Andrew Odgers, Managing Director
Reading time: 13 minutes
The problem, set out

Why This Needs Its Own List

Buying ordinary search work is difficult. Buying this is harder, for four specific reasons that combine badly. Our agency guides cover general vetting, so everything here is particular to this service.

No standards. Nothing defines what the service is.

There is no agreed scope, so two proposals under the same heading can describe entirely different work. A buyer comparing them is not comparing like with like and has no reference to check either against.

No certification. Nothing to hold anybody to.

No qualification, register or body exists for this. Anybody can describe themselves as doing it from tomorrow. A great many have.

No metric a buyer can verify. The most serious of the four.

In ordinary search a client can check claims against a console they own. Here there is no position, no citation report and no vendor dashboard, so a supplier's account of performance frequently cannot be tested at all. Our measurement guide sets out why.

A label new enough that nobody can correct its misuse. The one that ties the others together.

When a term is unfamiliar, a buyer cannot tell whether it is being used properly. That is a temporary condition in any field. While it lasts it is worth a great deal of money to whoever exploits it.

What follows for you. Judge process rather than claims.

Since outcomes cannot be verified here, the only thing worth assessing is how a supplier thinks and what they refuse to say. The rest of this page is about that.

This page is reviewed quarterly. The vetting logic outlasts the products.

The real workstreams

What An Agency Actually Does

Stated plainly, including the part that reduces what anybody can charge: most of this work is ordinary search optimisation. A buyer who knows that is much harder to overcharge.

Access work. The crawlers belonging to these products, checked and maintained.

Permitted at every layer, confirmed arriving in server logs, re-checked after any hosting or security change. This is genuinely additional to ordinary search work and it is the only part with documented consequences attached.

Content restructuring. Passages written so they stand alone.

The largest piece of work for most sites. Partly additional, since ordinary optimisation does not require it. Partly shared, since it improves the page for readers too.

Entity and record work. The business described identically everywhere.

Not additional at all. This was local search work long before any of this and it appears in our local SEO guides for that reason.

Earned references. Independent sources describing you.

Also not additional. Same discipline, same methods, same position on buying links that our backlinks guides have always taken.

Observation and reporting. Sampling, logs and a baseline.

Genuinely additional, since conventional reporting does not cover any of it.

What that adds up to. Three additional items out of five.

Access, restructuring and observation. If a proposal presents all five as new activity. You already pay somebody for the other two, you are being charged twice.

Hold two proposals against this

What A Service Should Include

Eight items, each with the reason it matters. A proposal missing three or more is narrower than it looks.

Crawler access verified in logs. Because it is the only documented lever and it fails silently.

Technical accessibility at every layer. Because a firewall can refuse what the robots file permits.

Content restructured for standalone passages. Because material that only makes sense in sequence is harder to use.

Entity consistency across the web. Because contradictions give an assembling system conflicting material about you.

Structured data matching the visible page. Because accuracy is the defensible reason to do it. It must never describe anything untrue.

Earned references. Because corroboration from independent sources is what a system can actually observe.

A recorded baseline. Because without a dated starting point nothing later can be compared to anything.

Sampling with the question set shown to you. Because you should be able to judge whether the questions are the ones your customers ask.

What should not be included. Three things.

A submission to any assistant, since none exists. A dedicated file or markup type, which Google's documentation checked on 30 July 2026 states is unnecessary for its AI features. And any guaranteed appearance, which nobody is able to provide.

One question. It settles most of it

The Question That Sorts Them

Ask how they will measure it. If you only take one thing from this page, take this. It separates capability from marketing more reliably than anything else you can ask.

What a capable answer contains. Four features.

It names specific proxies rather than a single figure. It describes sampling, including how many questions and how often. It admits what cannot be measured, particularly that no citation report exists. And it distinguishes evidence you own, such as your server logs, from things they will assert.

What a weak answer contains. A number.

Not proxies but a single figure: a visibility score, a share of voice, a percentage of answers you appear in. Presented as a measurement, with no method offered and none available if you ask.

Why this question works so well. It cannot be bluffed for long.

A supplier who has genuinely engaged with this subject has hit the measurement problem within a week, because it is unavoidable. One who has not will either produce a score or become vague. There is no third response.

The follow up that confirms it. Ask what the number would look like if things were going badly.

A real measure can move downward. A score constructed to demonstrate value generally cannot. Asking usually reveals which you are being shown.

And ask it of us. The answer is in our measurement guide, in advance, including what we refuse to report.

Six, all answerable

Questions To Ask

Each has a right kind of answer rather than a right answer. What you are testing is whether the supplier has thought about it.

Which of your claims has a vendor published rather than being your own observation? The most revealing question in the subject.

A capable supplier separates the two immediately, because they have had to. Anybody treating everything as equally established has not read the documentation or has decided it does not matter.

How much of this overlaps with work I already pay for? The answer should be most of it.

Anybody claiming almost no overlap is either mistaken or charging twice, per block two.

What happens when the systems change? Because they will.

Look for a review cadence and a willingness to revise position. Anybody whose plan has no mechanism for being wrong has not planned for the actual conditions.

What would you do first, for what reason? Sequencing reveals judgement.

Access and a baseline first is the answer we would give. Content production first suggests a supplier selling what they like making.

What would make you tell me to stop? The question almost nobody is asked.

A supplier with no answer has no threshold at which they would decline your money, which tells you what the engagement is for.

Who should not buy this service? A supplier who cannot name anybody is describing a sales target rather than a market.

Five, in rough order of seriousness

Warning Signs

Any one of these is worth a conversation. Two together is usually enough to stop.

A guaranteed appearance, which nobody is able to provide. Disqualifying rather than ambitious.

Nobody controls whether a system names a business, so this is a claim about something outside the supplier's control. It is the clearest single signal available to you.

A proprietary visibility score presented as fact. A proxy dressed as a measurement.

The underlying sampling may be sound. The number implies a precision that the variability of these systems cannot support. The method is almost never available for inspection.

Claims about internal mechanics no vendor has published. The subtle one.

Confident accounts of how sources are weighted, named ranking factors with percentages attached. None of that has been published by anybody operating these products, so it was inferred at best.

Premium pricing for the same work under a new name. Compare activities rather than headings.

If the workstreams match what your previous supplier delivered and the fee has risen substantially, you are paying for vocabulary.

Anybody who has not mentioned the fundamentals. The quietest signal and a reliable one.

A proposal that says nothing about accessibility, entity consistency or content quality has skipped the foundation of the whole thing. Our page on whether this is worth it explains why that ordering matters.

Uncomfortable for us too

Judging Their Own Content

Read what a supplier publishes about this subject. It is the cheapest due diligence available and it tests exactly the thing that matters, which is whether they know what they do not know.

What to look for. One distinction.

Do they separate what a vendor has documented from what they have observed? A page that attributes claims to published documentation with dates, then labels the rest as observation, has been written by somebody being careful.

What its absence tells you. One of two things.

Either they do not know which of their claims is documented. Or they knew and decided it did not matter. Neither is a supplier you want holding your budget. You cannot tell which from the outside.

The specific things to check for. Four.

Ranking factors listed with percentages, which nobody has published. A statement that structured data improves AI visibility, which no vendor has confirmed. Adoption statistics with no source. And any claim about how sources are selected, presented without attribution.

Why this is uncomfortable for us. Because it applies here.

Every page in this cluster attributes its documented claims with a date and labels the rest as observation. If we had not done that, this block would be a liability rather than an argument. You should hold us to the same test you apply to anybody else.

The corollary worth knowing. Careful content reads as less confident.

A supplier who admits uncertainty will always sound weaker than one who does not. In this subject that is the wrong way round. Knowing it is most of what protects you.

What each arrangement rewards

Pricing Models

Three arrangements are common. Each creates different incentives. None is improper. It is worth knowing what each rewards before signing one.

A separate premium line. Rewards making the service look distinct.

Where this is billed separately, the supplier has a commercial interest in the work appearing as different as possible from ordinary search optimisation. Given the overlap in block two, that incentive points away from accuracy. It deserves scrutiny rather than refusal, since a genuinely additional scope can justify it.

Part of a retainer. Rewards doing whatever currently matters most.

Where it is included, the supplier has no reason to inflate this subject relative to anything else, which is why we price it this way. The weakness is that a client cannot see what proportion of effort went where, so it requires reporting that shows the work.

A fixed project. Rewards finishing.

Suits the one off parts, being an access audit, a baseline and an initial restructuring pass. It suits the accumulating work badly, since references and consistency build over quarters rather than completing.

What to be wary of in any model. Payment tied to a placement.

Any arrangement paying on appearance in an assistant assumes an outcome nobody controls, which means either the supplier does not understand the field or the trigger is written loosely enough to always fire.

The question that cuts across all three. What am I not already paying for?

Our cost page answers it for our own pricing.

Making the comparison usable

Comparing Two Proposals

Neither proposal can prove an outcome, so comparing promised results is comparing two pieces of fiction. Four things can be compared instead.

The activity list, item by item. Against block three.

Tick off what each includes. Differences in scope are usually larger than differences in price and much easier to establish.

What each says it will not do. The most informative comparison of the four.

A proposal containing refusals has been written by somebody who has thought about the limits. One containing only capabilities has been written to win.

The measurement answer. Per block four, side by side.

One will describe proxies and sampling and admit limits. One will produce a number. That comparison is usually decisive on its own.

What each would do in the first month. Sequencing reveals priorities.

Access and a baseline suggests somebody who has done this before. A content calendar suggests somebody selling production.

What to do if they look identical. Ask the two hardest questions.

What would make you tell me to stop, plus who should not buy this. Answers to those diverge sharply even when proposals do not. Per our approach page we publish ours rather than waiting to be asked.

Generative engine optimisation

Ask how they
will measure it.

A capable answer names proxies, describes sampling and admits what cannot be measured. A weak one produces a number with no method. That single question separates capability from marketing. You should ask it of us.

What is included every month:

Crawler access checks Technical accessibility Content restructuring Entity consistency Earned references Recorded baseline Question sampling Monthly reporting

£350 per month, one target area. No setup fee, nothing billed separately.

The full guide series

Twenty guides.
One subject.

The mechanism, all four platforms, what you can influence, whether it is worth it yet, what it costs, how to measure it and how we approach the work ourselves.

Questions people ask

Choosing a GEO Agency

Why is this harder to buy than ordinary SEO?
Four reasons combining badly. No agreed standards, so two proposals under one heading can describe different work. No certification, so anybody can adopt the label. No metric you can verify, since there is no position, no citation report and no vendor dashboard. And a term unfamiliar enough that you cannot tell whether it is being used properly. Judge process rather than claims.
What is the single best question to ask?
How will you measure it. A capable answer names specific proxies, describes sampling including how many questions and how often, admits what cannot be measured, then distinguishes evidence you own from things they assert. A weak one produces a score with no method. A supplier who has genuinely engaged with this has hit the measurement problem within a week, because it is unavoidable.
How much should overlap with our existing SEO?
Most of it. Of five workstreams, three are genuinely additional: crawler access work, content restructuring for standalone passages. The sampling and log work to observe any of it. Entity consistency and earned references are not additional at all. If a proposal presents all five as new activity and you already pay somebody for the last two, you are being charged twice.
What are the clearest warning signs?
A guaranteed appearance in any assistant, which is disqualifying since nobody controls it. A proprietary visibility score presented as fact. Confident claims about how sources are weighted, which no vendor has published. Premium pricing for the same work under a new name. And anybody who has not mentioned the fundamentals, which is the quietest signal and a reliable one.
How can we check a supplier before speaking to them?
Read what they publish on the subject and look for one distinction: do they separate what a vendor has documented from what they have observed? Its absence means they either do not know which of their claims is documented or knew and did not care. Look specifically for ranking factors with percentages, unsourced adoption statistics, then claims that structured data improves AI visibility, which no vendor has confirmed.
Should it be a separate line or inside a retainer?
Each rewards something different. A separate premium line gives the supplier an interest in the work looking as distinct as possible from ordinary search optimisation, which points away from accuracy given the overlap. A retainer removes that incentive but hides the proportion of effort, so it needs reporting that shows the work. Be wary of any model paying on appearance in an assistant.