Every AI visibility tool on the market will tell you your presence rate "in ChatGPT." Almost none of them measured ChatGPT. They called an API with web search enabled, which is a different system with a different prompt, different tools, different ranking and a different audience, and then labelled the result with the name of the consumer product because that is the name customers recognise.
The gap between those two things is not a rounding error, and pretending it does not exist is the single most consequential piece of dishonesty in this category. It is also completely fixable: name the surface you actually measured, store the configuration, and let the reader decide how much to generalise.
What actually differs
Put the two side by side and the list of differences is longer than most people expect, and every item on it can move a presence rate.
The system prompt. The consumer app runs with a substantial system prompt written by the vendor for a consumer audience: tone, safety, formatting, how much to hedge, whether to name commercial products at all. Your API call runs with whatever system prompt you set, which is almost certainly not that one. The tendency to name vendors is exactly the kind of behaviour a system prompt shapes.
Memory and personalisation. The consumer app knows things about the person using it. It may remember that they run a small ecommerce business, that they asked about checkout tooling last month, that they prefer short answers. Your API call knows nothing. Whether that makes your measurement better or worse depends on what you are trying to measure, but it certainly makes it different, and it is why running visibility checks by hand in your own logged-in account is close to useless.
The retrieval configuration. How many searches the model is allowed to run, whether it can follow links, how many results come back, whether there is a re-ranking step. The consumer product tunes these; your API call takes defaults, or whatever you set.
The surface itself. A consumer answer is rendered in a UI with source chips, follow-up suggestions and a specific citation treatment. An API response is text and a structured source list. The information may be equivalent; the thing a human sees is not.
Two systems that share a brand name
Neither column is 'the truth'. They are different systems, and a number computed on one is evidence about the other rather than a measurement of it.
| Property | Consumer app | API with web search | Effect on a presence rate |
|---|---|---|---|
| System prompt | Vendor-written, consumer-tuned, not disclosed | Yours, or minimal | Directly shapes whether products get named at all |
| Memory of the user | Often present | None | Personalised answers are not comparable between users |
| Retrieval settings | Tuned by the vendor, opaque | Set by you, or defaults | Changes which sources are available to cite |
| Reproducibility | Poor. You cannot pin a version | Good. Model id and tools are recorded | Only the second can be re-run and checked |
| What it represents | What a real person sees | What the model does with a clean question | The proxy question in one line |
Why we still measure the API
Given all that, the obvious question is why not measure the consumer product directly.
Because a number you cannot reproduce is not a measurement. The consumer app cannot be pinned to a version, cannot be run in a guaranteed-clean context at scale, and cannot be re-run six weeks later to check a disputed figure. Its answers are personalised, which means they are not the same object between two observers. Building a time series on it produces a chart whose movements you can never attribute to anything.
The API has the properties a measurement needs. It is reproducible, it is version-pinnable, it is clean by construction, and every run can carry the exact configuration that produced it. It is a proxy, and it is a proxy whose relationship to the thing you care about is stable enough to track over time, which is the actual requirement for a metric. You are almost never asking "what is my absolute presence rate today." You are asking "is it going up," and a consistent proxy answers that better than an inconsistent direct measure.
The dishonesty is not in using the proxy. It is in not saying so.
"ChatGPT" is a false claim. "ChatGPT (API · web search)" is a true one, and it is not meaningfully harder to read.
The four surfaces, named honestly
- Perplexity (Sonar API): Perplexity's Sonar API.
- Google AI Overviews: the real AI Overview that Google renders on a live search results page.
- Claude (API · web search): the Anthropic API with the web search tool enabled.
- ChatGPT (API · web search): the OpenAI API with web search enabled.
Three of those are APIs. One is not, and the difference is worth dwelling on because it changes what that row of your dashboard means.
Google AI Overviews is a rendered surface, not an API call. It is what a person actually sees at the top of a search results page. That makes it, in one specific sense, a more direct measurement than the other three: there is no proxy step between the observation and the human experience.
It is also a genuinely different kind of thing from a chat answer. An AI Overview sits above ten blue links, is triggered for some queries and not others, and answers a search query rather than a conversation. A prompt written as a natural question (which is the right shape for the chat surfaces) is not what people type into Google. So the same prompt is doing two different jobs across your engine set, and a per-engine breakdown is not optional decoration. It is the only way to read the number at all.
Do not average across engines without looking underneath
An aggregate presence rate across four engines mixes three chat APIs and one search surface. If your aggregate moved, check which engine moved. A change concentrated in one engine is usually a change in that engine, not a change in you.
Store the configuration, or the number expires
Every stored answer carries the exact model identifier and tool configuration used to produce it. This is not archival tidiness; it is what stops the entire dataset from silently expiring.
Model versions change. They change without a changelog, without a version announcement, and sometimes without any external signal at all. A presence rate that fell 12 points between two scans has two candidate explanations (your visibility changed, or the model changed), and there is exactly one way to tell them apart: check whether the model identifier is the same.
Without that field, every movement in the series is uninterpretable. You will have a chart, and you will have a meeting about the chart, and nobody in the meeting will be able to rule out the possibility that the vendor shipped a new model on Tuesday.
With it, the first question in that meeting takes four seconds to answer, and half the time it ends the meeting.
Three things every run has to carry
Any one of these missing makes a movement in the series unattributable. All three are cheap to store and impossible to reconstruct later.
The model identifier
Which exact model answered. A silent version change is the single most common non-explanation for a metric moving.
The tool configuration
Whether web search was enabled and how it was configured. Two runs with different retrieval settings are not comparable.
The locale and timestamp
Answers vary by market and change over time. An undated observation is unusable within weeks.
What you can and cannot conclude
Being clear about the proxy is what lets you use it confidently, so here is the honest boundary.
You can conclude: that your presence in this configuration is rising or falling; that one engine treats you differently from another; that a specific prompt is one you lose; that a competitor is being named where you are not; that a page of yours is or is not being used as a source. All of these are comparisons within a consistent measurement, and comparisons within a consistent measurement are exactly what a proxy is good for.
You cannot conclude: that some specific percentage of real ChatGPT users see your name. Nobody can conclude that, with any tool, because no vendor reports impressions to publishers. There is no Search Console for AI answers, and a product claiming otherwise is either inferring it or making it up.
That is a smaller claim than the market generally makes, and it is a claim that survives being challenged. Which matters, because it will be challenged, usually by the person whose budget depends on the answer, and usually at the worst moment.
The same rule, everywhere else
This is not really a fact about answer engines. It is a general property of measurement that this field happens to violate loudly.
A number's label should name the thing that was measured, not the thing you wish you had measured. A crawl is not a citation. A session is not a person. A referral detected without a referrer header is an inference with a confidence attached, not an observation. In every one of those cases the shorter, wronger label is more marketable and the longer, truer one is what makes the number usable a year later.
The test is simple: could somebody reproduce this number from the label alone? If the label is "ChatGPT," no. If it is "ChatGPT (API · web search), model identifier recorded, locale recorded, clean context," yes, and reproducibility is the whole difference between a metric and a talking point.
Frequently asked
No. An API call with web search enabled differs from the consumer product in system prompt, memory and personalisation, retrieval configuration, and the interface the answer is rendered in. Every one of those can change whether a product gets named. The API is a reasonable proxy because it is reproducible and consistent over time, which is what a metric requires, but a number computed on it is evidence about the consumer product rather than a measurement of it.
Because consumer answers cannot be reproduced. They are personalised by account history and memory, cannot be pinned to a model version, and cannot be re-run in a guaranteed-clean context at scale. A time series built on them produces movements that can never be attributed to anything. The API surface gives up direct realism and buys reproducibility, which is the trade any usable metric has to make.
It is a rendered surface rather than an API call: the actual AI Overview that Google shows above the search results for a live query. That makes it closer to what a person sees, but it is also a different kind of thing from a chat answer. It responds to a search query rather than a conversation and appears for some queries and not others, so the same prompt does a different job there than on a chat API. Read the per-engine breakdown rather than the aggregate.
Because model versions change without a changelog, and a drop in your presence rate has two possible causes: your visibility changed, or the model did. The model identifier is the only way to tell them apart. Without it, every movement in the series is uninterpretable, and no amount of care in the statistics downstream can recover the distinction.
No. No assistant vendor reports impressions to publishers, so there is no equivalent of Search Console for AI answers and no way to observe real user sessions from the outside. What can be measured is how often a reproducible configuration names you, tracked consistently over time. Any product presenting an absolute share of real user exposure is inferring it, and the inference is not checkable.
Sources & further reading
- 01Web search tool, Anthropic
- 02Web search, OpenAI Platform
- 03Sonar API reference, Perplexity
- 04AI features and your website, Google Search Central