Every tool that measures AI visibility eventually gets asked for a single score. It is an entirely reasonable request. Executives want one number, dashboards want one gauge, and "your AI Visibility Score is 62" is a far easier sentence than the three-part answer it replaces.
It is also the point at which the measurement stops being usable. Not because averaging is imprecise, but because the three things being averaged have different causes, different fixes, and different populations underneath them. A score of 62 that fell to 54 tells you nothing about what to do, because there are at least three unrelated reasons it could have moved and the score deliberately discarded which one.
The three measures
Each is a fraction, and each fraction is worth reading slowly, because the numerator and the denominator both carry decisions.
Presence rate. Mentions divided by answered runs. Of the times we asked an engine your prompt and got an answer back, how often did that answer name you? This is the closest thing to "does the assistant know you exist in this context."
Citation rate. Runs citing a domain you own, divided by answered runs. Of those same answers, how often did the engine link to a page you control? This is the closest thing to "does the assistant use you as a source."
Share of voice. Your mentions as a proportion of every tracked brand's mentions. Not a rate against runs at all: a rate against your competitors. This is "of the airtime in this category, how much is yours."
The first two share a denominator and measure different numerators. The third has a completely different denominator and answers a completely different question. Averaging any two of them produces a number whose units do not exist.
What each number is actually a fraction of
Two share a denominator. One does not. This alone is why a blended score has no defensible definition.
| Measure | Numerator | Denominator | The question it answers |
|---|---|---|---|
| Presence rate | Runs that named you | Runs that returned an answer | Does the model know you belong in this conversation? |
| Citation rate | Runs that linked a domain you own | Runs that returned an answer | Does the model use your pages as evidence? |
| Share of voice | Your mentions | Every tracked brand's mentions | How much of the category's airtime is yours? |
| Prominence | Position-weighted mentions | Runs that named you | When you are named, are you named first? |
The four quadrants, and why the blend erases them
Presence and citation can each be high or low independently. That gives four states, and the whole diagnostic value of the measurement lives in telling them apart.
Named and cited. The healthy state. The model knows who you are and uses your pages to say so. Nothing to fix; protect it by keeping the cited pages fresh.
Named, not cited. The model has learned about you (from training data, from third-party coverage, from being talked about) and mentions you without linking anything you own. You get the credibility and none of the click. This is a retrieval problem, not an awareness problem: your pages either are not being crawled by the right agents, or are not extractable enough to be worth quoting. It is one of the few situations where the fix is genuinely about how the page is written rather than about what it says.
Cited, not named. Your page was good enough to use and your brand was not attached to the claim. The engine lifted a fact, credited a URL in a source list, and wrote a sentence that names somebody else. This is usually a page-structure problem: the passage that got extracted did not carry your name in it. Content that is anonymous in its useful paragraph gets used anonymously.
Neither. You are not in the conversation. Everything upstream comes first: crawl coverage, freshness, whether the prompts you are measuring are even the ones your buyers ask.
Now take the average. All four states can produce the same blended score, and the four fixes are: nothing, retrieval, page structure, and start over. A number that maps four different instructions onto one value is not a summary. It is a deletion.
You can be named without being cited, and cited without being named. Any score that averages the two is a number nobody can act on and nobody can check.
Prominence is a fourth thing, and it stays separate
There is an obvious refinement available: not all mentions are equal. Being the first vendor listed is worth more than being the fourth, and a position-weighted presence rate would capture that.
It exists, and it is deliberately kept as its own measure rather than being folded into presence.
The reason is that folding it in makes presence unfalsifiable. If "named first" scores higher than "named fourth" inside the same number, then a presence rate of 0.4 no longer tells you how often you were named. It tells you something about a mixture of frequency and position that cannot be decomposed back into either. You lose the ability to say "we were named in 8 of 20 answers," which is the one sentence in this whole area that a sceptical reader can check against the stored answers.
Keep them separate and both stay checkable: named in 8 of 20, and when named, usually second. Two facts, each verifiable, each pointing at a different piece of work.
One detail on position that matters for reading the numbers at all: many answers are prose, not ranked lists. There is no first and fourth in a paragraph. Those runs carry an explicit sentinel meaning "the answer had no ranking," rather than a position of zero or a null that some later average would quietly treat as worst place.
The denominators, where the real errors live
The interesting failures in this kind of measurement are almost never in the numerator. Counting mentions is mechanical. The errors are in what you divide by, and they are silent: the arithmetic is right, the number is plausible, and it is wrong by a factor nobody can see.
A brand's denominator is its own, not the site's. Each brand's rates are computed over the runs that brand was actually scored in, not over every run the site performed. Two ways this goes wrong if you use the site-wide total instead:
A day whose runs happened before your first brand was confirmed contributes answered runs that no brand was eligible to be scored against. Add those to a later brand's denominator and a brand that was named in four of its four runs renders as 0.267 across a wider window. Nothing looks broken; the number is simply a quarter of the truth.
And a competitor added mid-window has rows only from the day it was configured onward, while the site-wide denominator covers the whole window. That competitor's share of voice is then divided by runs it was never in, so a competitor you added yesterday appears to be doing badly, and the appearance is entirely an artefact of when you typed their name.
A scan is run once and scored against every brand. This is the one that gets products in trouble. The engine is called once per prompt and engine; that single answer is then evaluated against you and each of your competitors. If the underlying rows are keyed per brand, the run count and the cost are replicated across them, identically. A site tracking itself plus three competitors has four rows carrying the same cost, and summing them reports four times the real bill, with every test passing, because the arithmetic is correct and only the grain is wrong.
Any measure that is a property of the run (how many ran, how many answered, what they cost) must be counted once per prompt-engine-day and not once per brand. Any measure that is a property of a brand (mentions, citations) is summed per brand. Mixing the two is the single most reliable way to publish a confident wrong number in this area.
A question to ask any visibility vendor
"Is a competitor's presence rate divided by the runs that competitor was scored in, or by every run you performed for my site?" If it is the second, every competitor you added after the first day is understated, and the understatement grows the further back your date range reaches.
Zero, and the thing that is not zero
There is a state that looks exactly like 0% presence and is not: scans ran, cost money, and no brand was eligible to be scored against them.
This happens for real. A site whose only brands are unconfirmed automatic proposals has nothing confirmed to look for. The engine was asked, the answer came back, and no one had said what a hit would look like.
Reporting that as "0% presence" is a false sentence. It claims we looked and you were absent. The true sentence is that nobody has told us what to look for yet, and it has a completely different fix, which takes about thirty seconds and is not a content strategy.
So those runs carry an explicit marker rather than being counted as misses. They still count toward coverage and cost, because they really did happen and were really billed. They do not count as an absence, because they are not one.
This is the same distinction that governs every blank cell in the product: a zero is a measurement, and the absence of a measurement is not a zero.
Reading a run
Four outcomes, and only one of them is 'we looked and you were not there'. Collapsing the other three into it is how a visibility number becomes untrustworthy.
The call did not land
Not an answered run. It leaves the denominator entirely rather than counting as a miss.
Answered, no brand configured
Marked as unscored. Counts toward coverage and cost; does not count as an absence.
Answered, brand configured, not named
A real miss. This is the only thing that should lower a presence rate.
Answered and named
A hit, with a confidence score and the method that established it, never a bare boolean.
Every mention carries how it was found
One last property that follows from the same principle. A mention is never recorded as a bare yes. It carries a confidence between 0 and 1 and the method that established it:
alias_exact: the canonical brand name, matched on word boundaries.alias_variant: a configured alias (a product name, a common misspelling, a social handle).domain_reference: a domain you own named in the prose rather than in the source list.
The reason to keep the method rather than just the verdict is that these are not equally strong evidence, and a disputed number has to be traceable to something. "We were named 8 times" is an assertion. "We were named 8 times, of which 6 were exact matches on the brand name and 2 were the product name, and here are the sentences" is a claim somebody can check and disagree with specifically.
That is the same rule the detection engine has always followed: every classification records its method, because a customer who cannot audit a number has to either take it on faith or throw it away, and most people sensibly throw it away.
What to do with three numbers
Read them in this order, and let the pattern name the work.
If presence is low, nothing downstream matters yet. Either your prompts are not the ones your buyers ask, or you are genuinely not in the category conversation. Check the prompt set first: it is cheaper to be wrong about than the alternative.
If presence is fine and citation is low, you are known and not used. Look at crawl coverage for the pages that should be answering these questions, then at whether those pages contain a self-contained passage worth quoting.
If both are fine and share of voice is falling, you have not got worse. Somebody else has got better, and the useful artefact is the list of who is being named instead of you, which is a category map you did not draw and should probably read.
Frequently asked
Presence rate is how often an engine names your brand in its answer, as a fraction of the runs that returned an answer. Citation rate is how often that answer links to a page on a domain you own, over the same denominator. They come apart in both directions: a model can name you from what it learned without linking anything of yours, and it can use your page as a source while crediting somebody else in the prose. The two combinations point at different fixes, which is why they are never blended.
Because presence, citation and share of voice have different denominators and different causes. Named-but-not-cited is a retrieval problem, cited-but-not-named is a page-structure problem, and neither is an upstream crawl or prompt problem. All three can produce the same blended score, so the score maps several different instructions onto one value and discards which one applies. A single number is also unverifiable against the underlying answers, while 'named in 8 of 20 runs' can be checked directly.
Your brand's mentions as a proportion of every tracked brand's mentions across the same runs. Unlike presence and citation rate it is not measured against the number of runs at all, but against your competitors, so it answers how much of the category's airtime you hold rather than how often you appear. It can fall while your presence rate is flat, which means somebody else improved rather than you declining.
Check what denominator is being used. A brand added mid-window has scored runs only from the day it was configured, so dividing its mentions by every run performed for your site across the whole window understates it by exactly the fraction of the window it was absent for. A correct implementation scopes each brand's rates to the runs that brand was actually scored in, and nothing on the number's face reveals which one you are looking at.
Not necessarily. It can also mean scans ran but no confirmed brand was configured for them to look for, which is a completely different situation with a thirty-second fix. Those runs should be marked as unscored rather than counted as misses: they still consumed coverage and cost, but 'nobody told us what to look for' is not 'we looked and found nothing'. Check that your brand is confirmed before reading a zero as a verdict.
Sources & further reading
- 01GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (arXiv:2311.09735)
- 02Sonar API reference, Perplexity
- 03Web search tool, Anthropic
- 04AI features and your website, Google Search Central