Ask an answer engine the question your buyers ask, and it will hand you a list. The interesting thing on that list is rarely your own position. It is who else is on it.
That list is a category definition, and you did not write it. It was assembled by a model from whatever it read: your site, your competitors' sites, review platforms, forum threads, a Reddit comment from 2024, a listicle written by somebody who had used none of the products. Whatever produced it, it is now the answer a real buyer gets, and it is the closest thing that exists to an external, unflattering map of the market you are actually in.
Most teams never look at it, because their reporting shows their own rate and not the roster.
Two lists, and the space between them
There are two competitor sets in play and they are almost never the same.
The set you configured. The brands you told the product to look for. It comes from your positioning deck, your win-loss notes, and the two companies that come up on every sales call. It is a list of who you think you compete with.
The set that appears. Every brand actually named in the stored answers. It comes from whatever the model learned. It is a list of who a buyer is shown alongside you.
The first list produces your share-of-voice number. The second list is where the intelligence is, and the delta between them is the thing worth an hour of somebody senior's attention.
Three shapes of delta, each meaning something different:
A name you have never heard of appears repeatedly. Usually one of three things: a genuinely new entrant, an adjacent product the model considers substitutable, or a content operation with no product behind it that has out-published everybody. All three matter. The third one matters most, because it is the cheapest to counter and the most galling to lose to.
A competitor you fight constantly is absent. They are not in the model's version of the category. This is a real advantage and a short-lived one. It also means your head-to-head prompts against them are testing a comparison the model does not natively make, which is worth knowing before you read those numbers.
The set is dominated by a category label you would not use. The strongest and most uncomfortable signal, and the one the next section is about.
When the engine puts you in the wrong category
Sometimes the roster is not a list of competitors at all. It is a list of products from an adjacent category, and your presence on it means the model has filed you somewhere you did not intend.
This is not a measurement error. It is the finding.
An answer engine's implied category is an aggregation of how the whole web describes a space, weighted by whatever the retrieval step surfaced. If it puts you next to general web analytics tools when you sell AI-traffic attribution, that is evidence about how the market describes you, and the market includes every third-party article that ever mentioned you in a list.
You have three honest responses, and the choice between them is a strategy decision rather than a marketing one.
Accept the framing. If the adjacent category is where the demand is, being a default answer inside it may be worth more than being the only answer in a category nobody searches for yet. Uncomfortable, occasionally correct.
Contest it. Publish the distinction, repeatedly and in the shape retrieval likes: a self-contained passage that states plainly what you are and what you are not. Category-defining content is one of the few genuinely effective moves here, because the model is summarising a corpus and you can add to the corpus.
Reframe the prompts. If your buyers do not use the category label you built the prompt set around, your prompts are testing a vocabulary nobody types. Check this before concluding anything, because it invalidates the measurement rather than reporting on it. It is the first thing to verify when the numbers look wrong.
The failure that looks like a competitor problem
If the roster is full of companies from a category you do not belong to, do not immediately conclude you are losing to them. First check whether your prompts described that other category. A prompt written in the wrong vocabulary retrieves the wrong industry and then reports faithfully on it.
A proposal is not a measurement
Brands can be proposed automatically from what the answers contain. That is how a name you never heard of ends up in front of you at all.
A proposed brand is not scored until it is confirmed. This looks like friction and it is a correctness rule.
An automatic proposal is a guess: a capitalised string that appeared near a product discussion. Some of those guesses are companies. Some are product features, open-source libraries, a person's surname, or your own brand spelled differently. Scoring them all would produce a leaderboard full of things that are not competitors and a share-of-voice denominator inflated by noise, and the denominator is exactly where visibility arithmetic goes wrong quietly.
So the confirmation step is where a human says "yes, that is a company, and yes, it belongs in this comparison." Before that, it is a suggestion. After it, it is part of a measurement.
Two consequences worth internalising:
Confirm early or your history is thin. A brand confirmed today starts being scored today. Its rates cover only the runs it was actually in, which is correct and also means you cannot retroactively learn what its share of voice was last quarter. The roster is worth reviewing weekly for this reason alone.
A brand you deliberately did not confirm still appears in the answers. It is just not in the leaderboard arithmetic. The stored answers remain the complete record, which is where you go when somebody asks whether a particular company ever came up.
Reading a share-of-voice move
Share of voice and presence rate move independently. Which one moved tells you whether the change is about you.
| Presence rate | Share of voice | What actually happened | Where to look |
|---|---|---|---|
| Flat | Falling | A competitor improved. You did not decline | Which brand gained, and on which prompts |
| Falling | Flat | The whole category is being named less, or answers got shorter | Answer length and structure in the stored runs |
| Falling | Falling | You specifically lost ground | Crawl coverage first, then the cited pages |
| Rising | Flat | The category expanded and you grew with it | New brands on the roster: the set is widening |
The first row is the one most often misread. A falling share of voice with a flat presence rate is not a decline, and treating it as one sends a team off to fix pages that are working. Somebody else got better, and the useful artefact is their name plus the prompts they won.
Read it per prompt, not in aggregate
An aggregate roster across your whole prompt set is a blur, because different questions produce genuinely different categories.
"Best analytics for AI traffic" and "how do I prove AI search is driving revenue" will return overlapping but distinct sets: the first a product category, the second a mix of products, agencies and consultancies, because a problem-led question does not presuppose that the answer is software. Averaging those into one leaderboard produces a category that exists in no answer anyone received.
The per-prompt view is the one to actually work with. For each prompt: how you did, who was named more often, and which pages were cited instead of yours. That last column is the most directly actionable thing in the entire product, because it is a list of specific URLs that beat a specific page of yours on a specific question. You can open them.
Not "are we winning." Which question are we losing, to whom, and with which page.
What to do with the roster
Every week: scan for new names. New entrants appear in answers before they appear in your win-loss data, because a model reflects published material and your sales team reflects deals that reached a call. This is genuinely earlier intelligence than the pipeline gives you.
Every month: read the pages that beat you. Take the three prompts you lose most consistently, open the URLs cited instead of yours, and read them as a retriever would: does one passage answer the question completely, on its own, with the vendor named in it? The answer is usually yes, and it is usually the only difference.
Every quarter: check whether the category is drifting. The set of brands named beside you changes as the corpus changes. A slow drift toward an adjacent category is a strategic signal, and it is only visible if somebody is looking at the roster rather than at the rate.
And once, now: compare the set you configured against the set that appears, and be honest about which one your positioning was written for.
Two lists
The brands you configured and the brands that actually appear. The gap between them is the finding, not an error
review it weekly
Confirm first
An automatically proposed brand is a guess until a human confirms it. Unconfirmed proposals are never scored
a proposal is not evidence
Per prompt
Different questions produce different categories. An aggregate roster describes a category that appears in no answer
aggregate to report, per-prompt to act
Frequently asked
Your brand's mentions as a proportion of every tracked brand's mentions across the same engine runs. It is measured against your competitors rather than against the number of runs, so it answers how much of the category's airtime you hold rather than how often you appear. It moves independently of presence rate: a falling share of voice with a flat presence rate means a competitor improved, not that you declined.
Because the category an engine describes is aggregated from everything it read, not from your positioning. A name you do not recognise is usually a new entrant, an adjacent product the model treats as substitutable, or a content operation that has out-published the actual vendors. All three are worth knowing about, and the third is both the cheapest to counter and the easiest to miss, because it will never appear in your win-loss data.
First check whether your prompts described that other category, because a prompt written in the wrong vocabulary retrieves the wrong industry and then reports on it accurately. If the prompts are right, you have three options: accept the framing if that adjacent category is where the demand actually is, contest it by publishing self-contained passages that state what you are and are not, or reframe your prompt set if your buyers do not use the label you built it around.
Because an automatically proposed brand is a guess extracted from answer text, and some of those guesses are product features, libraries, surnames or your own brand spelled differently. Scoring them all would fill the leaderboard with non-competitors and inflate the share-of-voice denominator with noise. Confirmation is where a person decides the name is a company that belongs in the comparison. Unconfirmed proposals still appear in the stored answers, they are just not part of the arithmetic.
Per prompt, when you intend to act on it. Different questions produce different categories: a product-category question returns vendors, while a problem-led question can return a mix of products, agencies and consultancies, since it does not presuppose the answer is software. Averaging those into one leaderboard describes a category that appears in no answer anyone actually received. Use the aggregate for reporting and the per-prompt view for work.
Sources & further reading
- 01GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (arXiv:2311.09735)
- 02Sonar API reference, Perplexity
- 03AI features and your website, Google Search Central
- 04Web search tool, Anthropic