There is an enormous amount of advice about how to get cited by AI assistants, and almost none about how to tell whether any of it worked. This is a strange state for a discipline to be in. Search engine optimisation had rank tracking within about eighteen months of becoming a job. Generative engine optimisation is several years in and most practitioners still cannot answer "did last quarter's work do anything."
The reason is structural rather than lazy: there is no console. No platform gives you impressions, positions, or citation counts. Everything has to be constructed from the outside.
This is a measurement framework for doing that — three layers, in dependency order, with the failure mode each one is designed to catch.
Why there is no console
Search engines had an incentive to give you data: you optimised for them, sent them better content, and their results improved. Search Console is a cooperation mechanism.
Assistants have no equivalent loop. They are not sending you traffic as their business model — they are answering a question, and the link is a courtesy or a citation obligation. There is no product reason to build you a reporting surface, and no sign that one is coming.
So GEO measurement is inherently external. You are studying a system you cannot instrument, from the outside, by sampling its outputs and watching your own edge. That constrains what is knowable, and being clear-eyed about the constraint is what separates a real measurement programme from a dashboard of comforting numbers.
Three layers, in dependency order
Each layer depends on the one before it. Movement in a later layer without movement in an earlier one means something other than your GEO work caused it.
Visibility
Are you mentioned or cited when assistants answer questions in your category?
Referral
Do those citations produce actual clicks arriving at your site?
Revenue
Do those visits convert, and what are they worth?
The trap
Reporting one layer while claiming another. Visibility is not traffic; traffic is not revenue.
Layer one: visibility
The question: when someone asks an assistant a question in your category, are you mentioned?
How to measure it: synthetic sampling. Build a prompt set, run it on a schedule against each assistant, and record what comes back. Nothing else is available.
The design of the prompt set determines everything, and most teams get it wrong in the same three ways.
Too few prompts. Assistant outputs are stochastic. The same prompt returns different answers on different runs, so a single-prompt check measures sampling noise. You need enough prompts, run often enough, that you are estimating a rate rather than observing an event.
Prompts phrased like keywords. "AI attribution tool" is a search query. Nobody types that into an assistant. They type "we can't tell how much revenue comes from ChatGPT, what should we use?" Prompt sets built by importing a keyword list measure a behaviour that does not occur.
Only high-intent prompts. Category-defining questions ("what is X", "how do I measure Y") matter as much as purchase-intent ones, because that is where an assistant forms its picture of who the credible players are.
Track three separate things per prompt, per assistant:
- Mention rate — how often you appear at all.
- Citation rate — how often you appear as a linked source, which is different and rarer.
- Position — where in the answer, since first-named carries disproportionate weight.
Measure each assistant separately, always
Published analyses of citation overlap have found only around 11% of domains cited by ChatGPT are also cited by Perplexity. These systems have different retrieval architectures, different indexes and different source preferences. A blended "AI visibility score" averages away the only actionable information in the dataset — which assistant you are failing with, and therefore what to do about it.
Layer two: referral
The question: do citations produce visits?
This is where GEO measurement meets the attribution problem, and where most programmes quietly stop being measurable. Most AI-referred visits arrive with no referrer, so the obvious approach — count visits whose referrer is an assistant — captures a minority, and a biased one.
What to track:
Detected AI sessions, by assistant, with confidence. Not a single number. A count per assistant with a confidence distribution, so you can distinguish growth from detection changes.
Landing page distribution. Which pages assistants send people to. This is the highest-value diagnostic in the entire framework, because it is the only direct feedback loop you have between content and outcome. If your comparison page gets cited and your pricing page does not, that is an instruction.
The unattributable share. Your Direct traffic on deep pages, tracked as a trend. It is a proxy, not a measurement, but its direction is informative, and pretending you can see everything is worse than tracking a proxy honestly.
Visibility does not imply referral
The two layers decouple more than people expect. An assistant can name you prominently and send no traffic at all, because the user got their answer and had no reason to click. This is not a failure of your GEO work — it may be brand exposure worth having — but it is not traffic, and reporting a mention rate as though it were traffic is the most common dishonesty in this field.
Layer three: revenue
The question: are these visits worth anything?
Same machinery as any other channel — with two AI-specific complications.
The window has to be long. The AI touch is a discovery touch. Short attribution windows exclude it structurally, and a 24-hour or single-session window will report that the channel produces nothing.
The model has to not be last-touch. Last-touch erases this channel by construction, because the final touch before purchase is almost always a branded search or direct visit. Measuring a discovery channel with a conversion-moment model produces a zero, and a zero is very persuasive to a finance team.
Report both an introduction view — revenue where an AI touch appears anywhere in the history — and a last-touch view, side by side. The gap between them is itself the finding.
What to actually put on a dashboard
A minimal GEO scorecard
Per assistant, tracked over time. Absolute values matter less than direction.
| Layer | Metric | Source | Watch for |
|---|---|---|---|
| Visibility | Mention rate across the prompt set | Synthetic sampling | Noise from too small a sample |
| Visibility | Citation rate (linked, not just named) | Synthetic sampling | Rarer than mentions; do not conflate |
| Referral | Detected AI sessions with confidence | Your own edge | Detection changes masquerading as growth |
| Referral | Landing page distribution | Your own edge | The most actionable metric here |
| Referral | Deep-page Direct traffic trend | Your own analytics | A proxy — label it as one |
| Revenue | Introduction-view revenue | Attribution pipeline | Requires a long window |
| Revenue | Conversion rate vs non-branded organic | Attribution pipeline | The comparison that matters |
What the evidence says actually moves visibility
Setting aside measurement for a moment: the published research on what correlates with citation is more consistent than the field's noise level suggests.
+115%
Reported visibility lift from proper outbound citation, for mid-authority sites
Largest single effect
+37%
Reported lift from including expert quotations
Attributable statements
+22%
Reported lift from including statistics
Specific, sourced figures
Treat these as directional rather than precise — they come from vendor analyses with the methodological caveats that implies, and the same scepticism applied to conversion benchmarks applies here. But the pattern is coherent and mechanically sensible: content that is extractable, specific and corroborated gets cited more than content that is vague and unsourced.
That is not a growth hack. It is the observation that a system built to synthesise reliable answers prefers sources that behave like reliable answers. The tactic and the honest version of the work happen to be the same thing, which is a pleasant property and not one you should expect to last if it stops being true.
Two further findings worth internalising:
Third-party mentions carry weight. Discussion forums and professional networks appear disproportionately in citation analyses. Assistants weight corroboration across sources, and being described by others is a different signal from describing yourself.
Freshness matters differently per assistant. Retrieval-heavy assistants favour recently updated pages; others favour established, well-corroborated references. This is another reason the per-assistant breakdown is not optional — the same page can be optimal for one and mediocre for another.
The discipline
Two rules that prevent most self-deception:
Never claim a layer you did not measure. If you measured mention rate, report mention rate. Do not report it as visibility, do not imply traffic, and do not let it appear on a slide next to a revenue figure without a very clear separation between them.
Establish the baseline before you start. The most common failure in GEO measurement is beginning to measure after the work has already begun, leaving no counterfactual. Run the prompt set for a month before changing anything. That month is worth more than the next six of optimisation, because it is the only thing that lets you attribute a change to a cause.
Frequently asked
In three layers. Visibility, measured by running a prompt set on a schedule against each assistant and recording mention rate, citation rate and position. Referral, measured at your own edge by detecting AI-referred sessions and their landing pages. Revenue, measured by attributing those sessions with a long window and a model that is not last-touch. Each layer needs its own metric, and claiming one while measuring another is the standard mistake.
No, and there is no indication one is coming. Search engines had a business reason to give site owners data; assistants do not have the same loop. All GEO measurement is external — synthetic sampling of assistant outputs, plus whatever you can observe from requests arriving at your own server.
Enough that you are estimating a rate rather than observing single events, because assistant outputs are stochastic and the same prompt returns different answers across runs. Cover the full funnel — category-defining questions as well as purchase-intent ones — and phrase them the way people actually talk to assistants rather than importing a keyword list.
It is a bad idea. Published analyses have found only around 11% of domains cited by ChatGPT are also cited by Perplexity, because these systems have different retrieval architectures and source preferences. A blended score averages away the only actionable information you have: which assistant you are failing with.
That is a real and common outcome. An assistant can name you prominently while the user gets their answer and never clicks. It may still be worth having as brand exposure, but it is not traffic and should not be reported as traffic. It is also the clearest evidence that the two layers must be measured separately.
Sources & further reading