AI visibility: how to measure it
What AI visibility means, what can honestly be measured, and how to read the numbers without over-reading them.
Ask ChatGPT the same question twice and you can get two different answers. Ask it from two accounts and you can get three. This is the central problem with measuring AI visibility, and most of the numbers being reported in this field are produced by ignoring it.
A measurement you can act on has to survive that variance rather than pretend it away. This guide is about how to build one — and, as much, about what a number of this kind can and cannot support once you have it.
What is being measured
“AI visibility” is not one thing, and collapsing it into a single percentage is where most reporting goes wrong. There are three distinct outcomes, and they mean different things to a business:
- Mention. The answer names you. The weakest outcome, and the most common.
- Recommendation. The answer names you as an option in response to a question about what to use. Meaningfully stronger, and the one that maps to intent.
- Citation. The answer links to a page of yours as a source. The only one that can send a person to your site, and the rarest.
A brand can be mentioned constantly and cited never — that is the ordinary state for a well-known company with a poorly structured site. The reverse also happens, and means something quite different. Any measurement that reports one number for all three has thrown away the distinction that decides what you do about it.
Position matters too, and is easy to over-read. Being named third in a list of five is not a rank in the sense a search position is a rank — the ordering is a property of one generated sentence, not of an index — but being named at all versus not being named is real.
The distinction that matters most
An answer produced with live web search and an answer produced from a model’s memory are not the same kind of evidence, and the difference is larger than any other methodological choice here.
A grounded answer — the model searched, retrieved pages, and composed from them — is evidence about your visibility today. You can act on it. A recalled answer, composed from what the model absorbed in training, is evidence about what the web looked like at some point before the training cutoff. It tells you something about brand presence historically. It tells you nothing about whether a change you made last month worked.
Filing a recalled answer as a measurement is the single most common error in this field, and it flatters incumbents systematically: an established brand is well represented in training data and will be named from memory regardless of what its site does now. A tool that does not record which of the two it got cannot tell you whether your numbers are about the present.
So the first question to ask of any AI visibility report is: for each answer, did the model search? If the report cannot say, the numbers are not measurements of the thing you think they are.
Prompts are the instrument
You cannot measure visibility in general. You measure it against a set of questions, and the set is the measurement — change it and every number changes with it.
Which puts real weight on how it is built:
- Questions your customers actually ask, in their words, not keyword strings. “best invoicing software for contractors” is a keyword; “what should I use to invoice clients if I’m a one-person contractor?” is what someone types into an assistant.
- A mix of intent. Category questions (“what tools do X”), comparison questions (“X vs Y”), problem questions (“how do I do X”) and brand questions (“is X any good”) behave completely differently. A set of only one kind measures one narrow thing and reports it as visibility.
- No prompt that names you unless it is meant to. Asking “what is Acme?” and recording that Acme was mentioned measures nothing at all. It is the most common way a visibility score gets inflated.
- Visible and editable. If you cannot see the questions a score was computed from, you cannot judge the score. This is the reason CiteSite shows every prompt and lets you change it.
- Stable over time. A prompt set that changes between runs makes the trend meaningless, because the instrument moved. Add prompts by all means; keep the originals and track them separately.
Variance, and what to do about it
Generated answers are not deterministic. The same prompt, the same engine, the same day can produce a mention on one run and nothing on the next. That is a property of the systems, not a bug in the measurement, and there are only three honest responses to it.
Repeat. One answer is an anecdote. A prompt asked across a set of engines on a schedule accumulates into something a trend can be read from.
Band rather than point. A four-point move in a score built on variable answers is noise, and an interface that animates it is manufacturing a story. Report a range or a band; the direction over weeks is the signal.
Never explain a single flip. The temptation to account for one answer changing is enormous and almost always produces a wrong explanation, which then becomes a plan. If it did not persist across runs, it did not happen.
Engines are not interchangeable
Averaging across engines hides more than it shows. They retrieve differently, cite at different rates, and are used for different things. Perplexity surfaces sources prominently and cites heavily; ChatGPT composes more and links less; Gemini draws on Google’s index and is subject to a separate robots.txt switch that most sites have never looked at.
Report them separately. A single blended figure moving down is not actionable, and half the time it is one engine changing behaviour rather than anything about you.
Cost pushes the other way and is worth saying out loud, because it shapes what any product can offer: a grounded answer is not free, and the price varies enormously between providers. That is why sensible measurement runs the cheap engine weekly and the expensive one monthly rather than everything every day — and why a tool offering unlimited real-time monitoring at a low price is worth asking hard questions about.
What a number like this cannot support
Attribution to a specific change. You added schema, and three weeks later mentions rose. Those two facts are compatible with your change working, with a competitor’s page being deindexed, and with the model being updated. Nothing available to you distinguishes them. Directional evidence over months, across many prompts, is the strongest claim the data supports.
Comparison to somebody else’s score. Different prompt set, different instrument, different number. These scores are internally comparable over time and not comparable between tools.
Traffic forecasting. Answer engines frequently resolve a question without a click. Being cited more does not straightforwardly become more visits, and any model converting one into the other is guessing.
Building a measurement you can trust
- Write 20–40 prompts your customers would actually type, mixed across the four intents, none of them naming you unnecessarily.
- Ask each engine directly, with web search enabled where the model supports it.
- Record for every answer whether it was grounded or recalled, and never let a recalled answer into a measurement of the present.
- Record mention, recommendation and citation separately, along with which competitors appeared.
- Repeat on a schedule, keep the prompt set stable, and read the direction rather than the point.
- Keep the raw answers. When a number moves, the answer text is the only thing that explains it.
Before any of it, check the site can be read at all — a measurement of an unreachable site measures the block, not the brand. The visibility checker covers the five preconditions in order, and the GEO guide covers the work behind them.
This is what CiteSite does: a visible, editable prompt set, asked across ChatGPT, Gemini and Perplexity on a schedule, with mention, recommendation and citation recorded separately — and the provenance of every single answer carried into every export.