Everything you will ever learn about your AI visibility is downstream of the questions you chose to ask. Get the prompt set wrong and every number after it is a precise measurement of the wrong thing, and unlike most measurement errors, this one is invisible. The chart looks fine. The intervals are honest. The rates are computed correctly over a set of questions nobody asks.
Building the set is an hour of work, most of it thinking rather than typing, and it is the highest-leverage hour in the entire programme. This is how to spend it.
The four intents, and why the mix is the design
A presence rate is not one quantity. It means something different depending on what was asked, and a set that is 80% one kind of question produces a number that describes 80% one thing.
Category discovery. "Best CRM for small teams." No brand named, widest funnel, hardest to win. This is where new demand is decided, and where your presence rate will be lowest and most volatile. It is the number that matters most and moves least.
Head-to-head. "Acme vs Globex." Two named brands compared. High commercial intent, and the one place where being absent is unambiguously bad: the buyer has already shortlisted, and you are either on it or you are not.
Problem-led. "How do I stop losing checkout sessions." A problem, no category named. The buyer has not yet decided what kind of product solves this. Presence here means the engine associates you with a problem rather than a category label, which is often the earliest and most durable form of visibility you can have.
Brand-direct. "What does Acme cost." Your own brand named. You will appear. It measures almost nothing about demand.
A set is worth trusting when all four are represented and you know roughly what proportion each holds. Weight it towards category discovery if you are trying to grow the top of the funnel; towards head-to-head if you are losing deals late. What you must not do is let the mix happen by accident, because the mix silently sets what your headline number means.
What a presence rate means, by intent
The same 30% presence rate is four different findings depending on which bucket produced it. This is why the mix has to be deliberate.
| Intent | Example | What presence means here | What absence means |
|---|---|---|---|
| Category discovery | best analytics tool for AI traffic | You are a default answer in your category | The category conversation happens without you |
| Head-to-head | Traceten vs Plausible | You are on the shortlist when a buyer compares | Someone asked about you specifically and got nothing |
| Problem-led | why does my AI traffic show as direct | You own the problem, not just the category label | A page that should rank for this does not exist or is not retrievable |
| Brand-direct | what does Traceten cost | Almost nothing. This is a control | Something is genuinely broken. Investigate immediately |
Brand-direct prompts are a control, not a score
This deserves its own section because it is the most common way a visibility programme flatters itself into uselessness.
If you fill your set with prompts naming your own brand, your presence rate will be excellent. It will also be a measurement of whether an answer engine can read your homepage, which is a question with a known answer.
Keep one or two anyway. Their value is entirely diagnostic: a brand-direct prompt that stops returning you is a genuine alarm, and it fires before anything else does. It means your pages have become unretrievable (blocked, moved, broken, de-indexed by whatever passes for an index), and it will show up here weeks before your category numbers move, for the same lag reason that everything in this area shows up late.
So: two brand-direct prompts as a smoke test. Not eight as a score.
How many prompts, and the arithmetic behind the answer
The instinct is to ask for as many as possible. The constraint is that every run is a paid API call, and your total run count is prompts × engines × repetitions × scans.
That multiplication is what actually determines whether your numbers mean anything, and it is worth doing before you write a single prompt. Twenty prompts across four engines at five repetitions is four hundred runs per scan, but those runs are spread across twenty different questions, so any individual prompt has twenty observations behind it, and twenty observations is a very wide interval.
That is the real trade, and it is not the one people expect:
More prompts gives you a better-covered category and a more trustworthy aggregate rate, at the cost of being unable to say much about any single prompt.
Fewer prompts, more repetitions gives you sharp per-prompt numbers about a narrow slice of the category, and an aggregate that generalises badly.
For most teams the aggregate is what gets reported and the per-prompt view is what gets investigated, which argues for breadth. But if there are two or three questions you genuinely care about winning (the head-to-head against your main competitor, say), those are worth their own concentrated attention rather than being averaged into a category number.
4 intents
Category discovery, head-to-head, problem-led, brand-direct. Cover all four, and know the proportions you chose
the mix defines the number
1 to 2
Brand-direct prompts. Enough to detect breakage, not enough to inflate your headline rate
a control, not a score
3 minimum
Repetitions per prompt per scan. Below this the interval is too wide for the result to be a measurement
engines are not deterministic
Writing the prompts
Six rules. The first is the one that gets broken.
Write what a buyer types, not what you wish they typed. The failure here is subtle and near-universal: you write prompts using your own positioning language, the engine finds your page because your page uses that language, and you conclude you have strong visibility. What you have measured is that your copy is internally consistent. If your category page says "AI traffic attribution platform" and your prompt says "AI traffic attribution platform," the test is circular. Ask a customer what they searched for before they found you, and use their words even when they are worse.
Use natural sentences, not keyword strings. People type questions into assistants, not queries. "best crm smb 2026" is a search-engine habit. "what CRM should a five-person sales team use" is what actually gets typed, and the two retrieve differently.
One question per prompt. A compound question returns a compound answer and the presence signal gets ambiguous: were you named in the part about pricing or the part about integrations? You cannot tell later, and the stored answer will show you a mention you cannot classify.
No dates, no "latest", no "currently." These bind your prompt to a moment and quietly change what it means as the moment passes. A prompt containing "2026" is a different question in March than in November, and the series is broken without anything looking wrong.
Match the locale to a market you actually sell in. The locale is part of the run and part of what makes it reproducible. Measuring your US visibility with prompts run in a locale where you do not operate produces numbers nobody should act on.
Avoid prompts whose answer is a single fact. "What year was Acme founded" has one right answer and no vendor list. There is no share of voice, no competitive information, nothing to learn. Prompts that invite a comparison are the ones that carry signal.
The circularity test
Before you commit a prompt, ask: would this sentence appear on our own website? If yes, rewrite it. A prompt built from your own vocabulary tests whether the engine can read your site, which is not the question you are paying to answer.
Freeze it, and mean it
Here is the discipline that makes the whole thing valid, and it is the one people abandon in month three.
The prompt set must not change. Every time you add, remove or reword a prompt, you break the series. The rate before the edit and the rate after it are measurements of different things, and no amount of care in the arithmetic downstream can repair that.
This is genuinely inconvenient. You will write a prompt, watch it for a month, and realise it was slightly wrong. The temptation to fix the wording and carry on is enormous, and it is exactly what destroys the comparison you spent a month accumulating.
The answer is not to never change anything. It is to start a second set and keep the first one running. Two series, one of which has more history, both of which are internally comparable. The cost is a few more runs; the alternative is a chart with an invisible discontinuity in it that somebody will eventually build a decision on.
Why a prompt's text is immutable
The product enforces this rather than merely advising it, and the reason is worth understanding because it generalises.
A prompt's text cannot be edited once created. Stored runs reference the prompt. If you could change the text in place, every historical run would silently start claiming to have asked the new question, including the runs whose stored answers are right there to be read. The evidence and the label would disagree, and the label would win, because the label is what the interface shows.
Changing a question is therefore modelled as what it actually is: archive the old prompt, create a new one. The old runs keep pointing at the words that produced them. The series break is visible instead of hidden, which is the entire point.
An instrument that can be edited in place makes every measurement taken before the edit unexplainable.
It is the same argument for keeping funnel definitions in version control: a measurement instrument that changes without a diff turns every historical number into a mystery.
Two conditions that will corrupt the results
Independent of the prompts themselves, and both easy to get wrong.
Personalisation and memory. Run the prompts in a clean context with no account history. Otherwise you are measuring what an assistant has learned about you (its operator) rather than what it tells a stranger. The distinction matters most for exactly the people most likely to run the test manually: the founder, who has been asking that assistant about their own company for a year.
Undated observations. Record the date, the model identifier and the tool configuration on every run. Answer engines change under you with no changelog and no version announcement. An observation without those three fields is unusable six weeks later, because you cannot tell whether a change in the number was a change in your visibility or a change in the thing measuring it. This is why the engine label is never a bare product name: "ChatGPT" is not a reproducible configuration.
A worked starting set
Twelve prompts, weighted towards discovery, for a hypothetical analytics product. Copy the shape, not the words.
- Category discovery (5). best analytics for tracking AI traffic · how to measure traffic from ChatGPT · tools that show which AI assistants send visitors · analytics that attributes revenue to AI search · what software tracks AI crawler activity
- Problem-led (3). why does traffic from AI assistants show as direct · how do I prove AI search is driving revenue · how to tell if AI crawlers are hitting my site
- Head-to-head (2). your product vs your closest competitor · your closest competitor alternatives
- Brand-direct (2). what does your product do · your product pricing
Note what is absent. No prompt contains a year. No prompt uses a phrase from the product's own marketing. No prompt asks two things. The head-to-head prompts name a real competitor, because that is what a real buyer would type.
Run it, leave it alone for six weeks, and resist every urge to improve it. The set you froze and doubted is worth more than the set you kept perfecting.
Frequently asked
Enough to cover your category, balanced across the four intent types, with the arithmetic done first: total runs equal prompts times engines times repetitions times scans, and each individual prompt only gets the runs allocated to it. Twenty prompts gives a trustworthy aggregate and thin per-prompt evidence; five prompts gives sharp per-prompt numbers about a narrow slice. Most teams report the aggregate and investigate per-prompt, which argues for breadth.
Only one or two, as a control. You will appear in answers about your own product, so a brand-named prompt measures whether an engine can read your homepage rather than whether there is demand for you. Their value is diagnostic: if a brand-direct prompt stops returning you, your pages have become unretrievable, and that alarm fires weeks before your category numbers move.
You can replace it, but not edit its text in place, and you should think of replacement as starting a new series rather than continuing the old one. Stored runs reference the prompt, so editing the text would retroactively change what past runs claim to have asked while their stored answers say otherwise. The right move when a prompt turns out to be wrong is to add a corrected one and keep the original running, so both series stay internally comparable.
Four things, in rough order of how often they occur. Prompts written in your own marketing vocabulary, which test whether an engine can read your site rather than whether buyers find you. Prompts containing a year or the word latest, which quietly change meaning over time. Compound questions, which produce mentions you cannot attribute to a part of the answer. And single-fact questions, which return no vendor list and therefore carry no competitive signal.
Because assistants personalise on account history and memory. Running your prompts in a session that has been discussing your own company for a year measures what the assistant has learned about you as its operator, not what it tells a prospective customer who has never heard of you. Use a clean context with no history, and record the date, model identifier and tool configuration on every run so the observation stays interpretable later.
Sources & further reading
- 01GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (arXiv:2311.09735)
- 02Sonar API reference, Perplexity
- 03Web search tool, Anthropic
- 04AI features and your website, Google Search Central