There is a version of AI visibility reporting that gives you one number and a trend line, and there is a version that tells you what to do on Monday. The difference between them is a single decision: whether being named in an answer and being cited as a source are tracked as one thing or two.
They are two. They come apart in both directions, constantly, and the four combinations are four different problems with four different owners. A team looking at a blended score is looking at a number that has already thrown away the part that would have told them which problem they have.
The four quadrants
Presence and citation are measured over the same runs and can each be high or low. Four states.
What each combination means and who fixes it
The fixes are unrelated to each other. This is the entire argument against a blended visibility score.
| State | What is happening | Most likely cause | Who owns the fix |
|---|---|---|---|
| Named and cited | The model knows you and uses your pages to say so | Working as intended | Nobody. Keep the cited pages fresh |
| Named, not cited | You are known from elsewhere; your own pages are not being retrieved | Crawl coverage, or passages that are not worth quoting | Whoever owns crawlability and page structure |
| Cited, not named | Your page was used and your name was not attached to the claim | The extracted passage does not contain your brand | Content. It is often a one-paragraph fix |
| Neither | You are not in this conversation at all | Upstream: crawl, or the prompt set is not what buyers ask | Start with the prompt set, then crawl coverage |
Named, not cited
The model produces your name. It does not link anything you own. Frequently it links a review site, a listicle, or a competitor's comparison page that mentions you.
What this tells you is that your brand exists in the model's parameters or in third-party coverage, and your own pages are not making it into the retrieval step. Two causes, and they are worth separating because the fixes are different.
You are not being crawled by the right agents. Training crawlers and search-index crawlers are different fleets with different purposes, and only the second makes you citable. A site that has been thoroughly crawled by training agents and lightly crawled by search-index ones will produce exactly this pattern: known, not retrievable. This is measurable directly, and it is the leading indicator that moves weeks before the traffic does.
Your pages are crawled and not worth quoting. The retrieval step selects passages, not pages. A page whose useful content is spread across six paragraphs, each of which requires the previous one for context, has nothing a retriever can lift out and use. The competitor whose page has a self-contained 60-word answer under a question-shaped heading gets cited instead, and their content is not better. It is more extractable.
The awkward part of this diagnosis is that it feels like a brand problem and it is a plumbing problem. Teams in this quadrant reliably reach for more content when the correct move is to make the existing content retrievable.
Cited, not named
This is the more surprising one and the one with the cheapest fix.
The engine linked your URL. It listed you as a source. And the sentence it wrote names somebody else, or nobody. You did the work, the model used it, and the credit went into a source chip that a substantial fraction of readers never click.
The mechanism is almost always the same. Retrieval works on passages. The passage that got selected (the crisp paragraph that actually answered the question) did not contain your brand name. The model had a good fact and no name attached to it, so it wrote a fact without a name.
The fix is to make the quotable passage self-identifying. Not by stuffing your brand into every sentence, which reads badly and gets less extractable rather than more, but by ensuring that the paragraph most likely to be lifted contains one natural occurrence of who is saying it. "Traceten hashes IP addresses at the edge under a per-site derived key" survives extraction with attribution intact. "IP addresses are hashed at the edge under a per-site derived key" does not.
Find these pages directly
The citations view lists the URLs engines linked to. Sort your own cited URLs and check, for each, whether the passage that answers the prompt contains your name. The ones that do not are your cited-not-named pages, and each is a paragraph edit.
There is a related signal worth knowing about. A mention can be established by a domain reference: your domain named in the prose of the answer rather than in the structured source list. That is a partial case of this quadrant: you got named, but as a URL rather than as a brand, which is better than nothing and worse than a name.
Joining citations to what happened next
A list of cited URLs is interesting. A list of cited URLs joined to what those pages then produced is a work queue, and the join is where the citations view stops being a vanity report.
Two columns do the work.
Answer fetches. Crawls made to answer somebody's live question, as opposed to crawls made to build a training set. This is the closest observable proxy for a real prompt actually touching your page: not an inference, an HTTP request that happened. A cited page with rising answer fetches is being used, right now, by people asking questions.
Revenue from AI landings. Revenue from sessions that landed on that cited page and arrived from an AI source. This is a deterministic join, not a model: a session landed here, it came from an AI source, it produced a payment. No attribution weighting, no decay curve, no assumptions. It is the narrowest and most defensible version of the number, which is exactly why it is the one worth putting in front of a finance team.
Together they turn "should we invest in AI visibility" into a table of pages sorted by money, and that is a fundamentally different conversation.
From an answer to a payment
Each hop is observable. None of them is inferred, which is what makes the resulting number survivable in a budget meeting.
An engine cites your URL
Recorded from the answer's source list, with the URL and its derived path stored separately.
Answer fetches hit that page
Server-side crawl records from agents fetching to answer a live question. An HTTP request, not an inference.
A visitor lands on it from an AI source
A session whose landing page is that URL and whose source is an assistant.
That session produces revenue
Joined through your payment system. Deterministic, and reported in USD.
Three things in that table you must not do
The join is honest about its own limits, and the limits are specific enough to be worth stating rather than hiding behind a tooltip.
Do not add up the answer-fetches column. The crawl rollup stores no hostname, so the match is on the page path only. If you serve several hostnames, two different pages that share a path will show the same count. The column is correct as a per-row indicator and wrong as a total.
Do not sum "peak engines in a day." There is no per-engine citation breakdown available, and that is a real limitation rather than a missing feature: the citation rollup records how many distinct engines cited a URL on its busiest day, not which ones. A distinct count on one day cannot be added to a distinct count on another day, because the same engine appears in both. Two days at "3 engines" is not six engines.
Do not read a blank cell as a zero. Where a figure cannot be measured for the range you selected, the cell is blank. A zero is a measurement: we looked and found none. A blank is the absence of one. The two look similar and mean opposite things, and this is exactly the distinction that keeps every other number on the page honest.
The blank you will hit most often has a mundane cause: session-level data has a shorter retention horizon than the visibility rollups. Pick a long enough date range and you get real citation counts sitting beside blank revenue, because the citations are still there and the sessions are not. Nothing is broken. The instrument simply has two different memories and says so.
Path only
How cited URLs join to answer fetches. The crawl rollup stores no hostname, so multi-hostname sites share a count
per-row yes, total no
Peak, not sum
The engines column is a per-day distinct count on the busiest day. It never adds across days
no per-engine breakdown exists
Blank ≠ 0
An unmeasurable figure renders blank. A zero means we looked and counted none
different sentences
Working the quadrants in order
If you do nothing else with this, do it in this sequence. It is ordered so that each step is only worth taking if the previous one came back clean.
1. Are you in the conversation at all? If presence and citation are both near zero, stop reading the visibility numbers and check the prompt set. The most common cause of total absence is measuring questions your buyers do not ask, and it is much cheaper to be wrong about than the alternative. Only after the prompts are right does absence mean what it appears to mean.
2. Named, not cited? Pull crawl coverage for the pages that should be answering these prompts. If search-index crawlers have not fetched them recently, that is the whole finding and the fix is access, not content. If they have, the problem is extractability and the fix is structural.
3. Cited, not named? Go to the cited page, find the paragraph that answers the prompt, and check whether your name is in it. This is the highest-return hour available in the entire discipline, and most teams never spend it because the quadrant is invisible on a blended score.
4. Named and cited? Read the revenue column, find the pages carrying it, and treat them as production infrastructure. They are cited because they are good and current; they stop being cited when they go stale, and nothing will alert you when that happens except this table.
The pattern that makes all four legible is the same one that makes six numbers better than one: a single figure at the end of a chain names no link in it. Two independent measures across the same runs name four states, and four states are a plan.
Frequently asked
It means the model knows you from training data or third-party coverage, and your own pages are not surviving the retrieval step. Two causes are worth separating. Search-index crawlers may not be fetching your pages, which is measurable from server-side crawl records and fixed by access rather than content. Or your pages are crawled but contain no self-contained passage worth quoting, in which case the fix is structural: make one paragraph answer the question completely, on its own.
Because retrieval selects passages rather than pages, and the passage that got selected did not contain your brand name. The model had a usable fact with no name attached, so it wrote the fact without the name. The fix is to make the most quotable paragraph self-identifying, with one natural occurrence of who is saying it, rather than writing the useful sentence in a voice that could belong to anyone.
Not per engine. The citation rollup records how many distinct engines cited a URL on its busiest day, which is why that column is labelled peak engines in a day. Because it is a per-day distinct count, it must not be summed across days: two days each showing three engines is not six engines. The per-prompt view is where you go for engine-level detail.
Most often because session-level data has a shorter retention horizon than the visibility rollups, so a long date range can show real citation counts beside a blank revenue cell. The cell is blank rather than zero on purpose: a zero would claim we measured revenue and found none, while a blank says the figure could not be measured for that range. Narrow the range and the revenue figure becomes available.
A crawl made to answer somebody's live question, as distinct from a crawl made to collect training data. It is the closest observable proxy for a real prompt touching your page, because it is an actual HTTP request rather than an inference. Rising answer fetches on a cited page indicate the page is being used to answer questions right now. It is only visible server-side, since crawlers do not run JavaScript.
Sources & further reading
- 01Overview of OpenAI crawlers and user agents, OpenAI
- 02PerplexityBot and Perplexity-User, Perplexity
- 03GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (arXiv:2311.09735)
- 04AI Insights: crawl-to-refer ratios and AI bot traffic, Cloudflare Radar