How to Measure Your Generative Engine Optimisation Performance
This is the unsolved part of the subject. There is no position to track, no universal report of citations and no vendor dashboard. Everything available is a proxy or a sample. Anybody selling you a single number has invented it.
There Is No Rank To Report
There is no position, no impression count and no universal report of when your material was used to answer somebody's question. That is the starting point of this page rather than a caveat buried at the end of it.
What does not exist. Four things, all of which people are sold. There is no ranking position inside an assistant. There is no count of how many times you were cited. There is no console the vendor provides for site owners. There is no agreed industry measure two suppliers would calculate alike.
Why we lead with this. Two reasons, one self interested. A client who understands the problem cannot be sold a fictional score by somebody else. They will also trust our reporting more, because we said what it cannot do before being asked. Both beat pretending otherwise.
What this page does offer. Four real methods with stated limits.
Sampling, server logs, proxy measures and the one genuine platform report that exists. None is a measurement of citation. Together they are considerably better than nothing. The difference between them and a score is that you can inspect how each was produced.
What we will not do. Convert them into a number. Every input here could be combined into an index and named. We will not, because it would look like a measurement while being an opinion with arithmetic attached.
This page is reviewed quarterly. The problem is stable even where the tooling is not.
Why Not
The distinction that matters here is that tracking a position is not merely hard. It is describing something the systems do not produce, which is a different problem and cannot be solved by better tooling.
Answers vary by design. The first reason.
Generated responses involve deliberate variability, so asking the same question twice can return different businesses with nothing having changed anywhere. A recorded position therefore describes one instance of a variable output.
Answers are personalised. The second.
Location, the surrounding conversation and in some products a stored memory of earlier exchanges all shape what comes back. Two people typing identical words are not running the same query, so whose result would a position even describe?
The systems underneath change. The third.
Model updates arrive on schedules nobody outside publishes in advance. A change on their side can move which businesses appear more than anything you did to your own site. That makes attribution genuinely hard rather than rhetorically hard.
There is no ordering to sample. The fourth. The decisive one.
Retrieval gathers a handful of sources and discards the rest, so there is no ranked list underneath from which a position could be read. Even where sources are displayed, the order tells you very little about what the answer leaned on. Our guide to how these systems assemble an answer covers that properly.
What follows. Better tools will not fix this. A more sophisticated tracker still samples a variable, personalised output with no underlying order. The problem is the object being measured rather than the instrument.
What Can Actually Be Measured
Real things, all of them partial. What matters is knowing precisely what each one does and does not establish, because the failure mode is treating a proxy as a measurement.
One genuine platform report exists. The most concrete item here.
Google's Search Central documentation, checked on 30 July 2026, states that sites appearing in its AI features are included in overall search traffic in Search Console, reported within the web search type. So Google traffic including these features is visible to you in a console you own. It does not separate AI feature clicks from ordinary ones, which is a real limitation rather than a quibble.
Referral traffic from assistant products. Where they identify themselves, visits arriving from them can be seen in analytics. This proves somebody clicked through. It says nothing about the far larger number of occasions you were named without anybody clicking.
Crawler activity in server logs. Evidence of retrieval rather than inference. Block five treats it separately, since it is the most underused method available and the only one producing hard evidence.
Branded search volume. A proxy for becoming known. Searches for your own name rising suggests the business is becoming a recognised entity. It cannot tell you why. Block six covers the limits of measures like this.
Manual sampling. The only method that observes the actual outcome.
Asking the questions and recording what appears. Crude, laborious and the only method that looks directly at whether you are named. Block four covers it.
Manual Sampling Done Properly
Done carelessly this produces anecdotes that mislead. Done properly it is the closest thing to evidence available. The discipline is what separates the two.
Fix the questions in advance. Then do not change them.
A set of questions your customers actually ask, written down before you start. Changing the set between rounds destroys the comparison, which is the commonest way this gets ruined.
Ask on a schedule. Monthly is enough.
The same questions, the same products, roughly the same day. Frequency matters less than consistency, since you are looking for direction over months.
Record more than whether you appeared. Four things.
Which businesses were named. Which sources were cited, where the product shows them. Whether your own site was the source or somebody else's page about you. And whether the answer was substantive at all, because in some categories nothing much happens.
Why that third item matters most. Being named because a directory listing you was cited is different from your own material being used. One is visibility you own. The other you borrow from a third party who may stop granting it.
Why one check tells you nothing. The variability from block two.
A single round is one instance of a variable output. Two rounds is a difference with no trend attached. Direction over three or more rounds is the first thing that means anything. It is still a sample rather than a measurement.
Control the conditions you can. Because the rest you cannot.
Signed out where the product allows it, plus consistent about location. That removes some personalisation. It does not remove the variability. No method does.
Server Logs Are Underused
Your own server records every request made to your site, including those from the crawlers belonging to these products. That is factual evidence rather than inference. Almost nobody looks at it.
What the logs actually contain. Verifiable requests.
Which crawler asked, which page it asked for, when, plus what your server returned. Vendors publish their crawler names and address ranges, so requests can be verified as genuinely theirs.
What it proves. Retrieval, which is the first stage of everything else.
If the crawler for a product is requesting your pages and receiving them, you are eligible to be drawn on by that product. If it is not, nothing else you do matters for that route.
What it does not prove. That you were used.
Being fetched is not being cited. The logs establish access rather than outcome. Treating crawler visits as a performance measure is a mistake we see suppliers make.
Why this is the item to insist on. It is yours. Logs come from your own infrastructure rather than from a vendor or a tool, so a supplier who has never looked has skipped the only hard evidence available.
The failure this catches that nothing else does. Silent refusal.
A site can permit a crawler in its robots file while a firewall or hosting layer refuses the requests. Nothing on the site looks wrong, no report shows a problem and the business simply never appears. The logs show it immediately, which is why we check them first.
Proxy Measures And Their Limits
These are the measures most often quoted as evidence of AI visibility. They are worth watching in aggregate. Individually each has an alternative explanation that is usually more likely.
Branded search volume. People searching your name.
Rising branded search does suggest growing recognition. It also rises from a leaflet drop, a van livery, local press, a recommendation or somebody else's advertising. Attributing it to AI visibility specifically requires ruling out everything else, which is rarely done.
Direct traffic. Visits with no referrer attached.
Frequently offered as proof that people heard about you somewhere unmeasurable. It is also where analytics puts traffic it cannot classify, including plenty that has nothing to do with any assistant. Treat a rise as a question rather than an answer.
Enquiry sources. Asking people how they found you.
Genuinely useful and systematically unreliable. People misremember, they compress several steps into one and they name the last thing they touched. Somebody told about you by an assistant, who then searched your name and clicked an ad, will tell you they found you on Google.
How to use them anyway. In combination, over quarters.
Several proxies moving together over several months is worth attention. One proxy moving in one month is noise. Presenting it as a result is how unfounded claims enter reporting.
The specific misattribution to watch. Your own name.
Reporting that mixes searches for your business name with searches for what you sell will show growth that is really just people who already knew you. Our comparison with ordinary search covers why that separation matters throughout.
What To Distrust
Three things. The reader who takes only this block away from the page has still got the valuable part of it.
Any tool presenting an AI visibility score as fact. The commonest one.
Such a tool has done something like the sampling in block four, then converted it into a number using a method you cannot inspect. The sampling may be sound. The number implies a precision that the underlying variability cannot support.
Any supplier reporting your position. There is no position.
A report showing you at a rank inside an assistant is describing something the system does not produce. This is not a matter of methodology. The thing being reported does not exist.
Any figure quoted without a method. Including favourable ones.
A percentage of answers you appear in, a share of voice, a citation count. Each needs a stated method, sample size and date to mean anything. Most are quoted with none of the three.
Four questions that settle it. Ask them of anybody, ourselves included.
Which products did you check, over how many questions? How many times did you ask each, given answers vary? What method turned that into this figure? And what would this number look like if the work were not going well?
What a capable answer sounds like. Less impressive than a bad one. It describes sampling, admits the limits and declines to give a single figure. Our guide to choosing a supplier treats this as the question that sorts capability from marketing.
What We Report
Stated so you can hold us to it, then compare it against whatever anybody else offers you.
Crawler access, from your logs. Confirmed rather than assumed.
Which crawlers reached your site, which pages, plus what your server returned. This is evidence. Its limit is that it proves access rather than citation.
A fixed sample of questions, monthly. With the set published to you.
You see the questions, so you can judge whether they are the ones your customers ask. Its limit is that it is a sample of a variable output rather than a measurement.
Search Console performance, including AI features. Because Google includes them.
Reported from the console you own rather than from a tool of ours. Its limit is that AI feature clicks are not separated from ordinary ones.
Branded against non branded search. Separated, always.
Searches for your name reported apart from searches for what you sell. Its limit is that neither is attributable to this work specifically.
Enquiries, counted. The only figure that pays anybody's wages.
What we do not report. Three things.
A position, because there is not one. A visibility score, because we would have to invent the method. And any figure we cannot show you the working for. Our approach page carries those as commitments rather than preferences.
The cadence. A written update every three weeks, a full technical audit quarterly, reporting monthly.
Setting A Baseline Now
Whatever method you settle on, recording a starting point today is the only thing that makes next year's comparison mean anything. This is the most valuable hour available in this subject and almost nobody spends it.
Why it matters more here than elsewhere. Because the measures are weak. Where measurement is precise a baseline is convenient. Where every measure is a proxy, the comparison against a starting point is doing most of the work.
What to record. Five things, in about an hour.
Which crawlers are currently reaching the site, from the logs. A first round of your fixed question set, with what appeared. Current Search Console performance. Branded and non branded search separated. And how many enquiries you had last month, counted properly.
Write down the date and the method. Not only the numbers. A baseline whose method nobody recorded cannot be repeated, which makes it useless the moment somebody asks whether the comparison is fair.
Why doing it late is worse than it sounds. You lose the year.
A business that starts measuring in month nine has an interesting figure with nothing to compare it to. No way to reconstruct what things looked like before. That is the whole cost of skipping this.
And it costs nothing. No tool, no subscription, no supplier required. Whatever you decide about the rest of this subject, this hour is worth spending before you decide it.
We will not invent
a number for you.
No position, because there is not one. No visibility score, because we would have to invent the method. No figure we cannot show you the working for. What you get instead is your own logs, a published question set and the console you already own.
What is included every month:
£350 per month, one target area. No setup fee, nothing billed separately.
Twenty guides.
One subject.
The mechanism, all four platforms, what you can influence, whether it is worth it yet, what it costs, how to measure it and how to judge a supplier selling it.