LLM Visibility & Optimisation
There is no rank tracker for what a language model says about you
Language models answer differently to different people on different days, from different sources. Some of that is measurable and some of it is not. This page sets out which is which, and what we do with the part that can be measured properly.
The honest starting point
Why classic measurement does not transfer
Rank tracking worked because search results were stable enough to sample once a day. Two people in the same city, searching the same phrase, saw close to the same ten links. A position was a real thing that could be recorded, charted and argued about.
None of that holds here. Ask an assistant the same question twice in a row and you can get a different set of businesses, a different framing and a different set of sources. Change three words in the prompt and the answer can change entirely. Ask from a different country, in a different session, with memory switched on, and it changes again.
So the first job is to stop pretending the old unit exists. There is no position. There are four possible results for any given question: you are cited with a link, you are named without one, you are described in some indirect way, or you are absent while somebody else is named. Those four are worth counting. A composite score sitting on top of them is not.
The second job is to be clear that a lot of what people want measured is not observable at all. Nobody outside the labs can see training data, nobody can see why a particular answer was produced, and no tool can tell you what an assistant will say tomorrow. We would rather say that in the first meeting than in a footnote on a quarterly report.
What changes
Rank tracking versus prompt-panel measurement
These are not two versions of the same instrument. They measure different things, need different sample sizes, and support very different conclusions.
| Dimension | Rank tracking in classic search | Prompt-panel measurement in assistants |
|---|---|---|
| The unit | A position for a keyword on a results page | A rate: how often you appear across repeated runs of a defined prompt |
| Repeatability | Two people in similar conditions usually see close to the same result | The same prompt can return different brands, sources and wording on consecutive runs |
| Sample needed | One check per keyword per day is meaningful | One run is an anecdote; a usable rate needs repeated runs in clean sessions |
| What the top result means | A defined slot with reasonably well understood click behaviour | There is no slot — only cited, named, described, or absent |
| Attribution to your work | Movement is usually traceable to a change you or a competitor made | Movement may be your work, a model update, a reworded prompt or chance, and the report has to say when it cannot separate them |
| Agreement between tools | Different rank trackers broadly agree with each other | Different visibility tools frequently disagree, because their prompts, models and run counts differ |
| Most useful finding | Where you sit against competitors for commercial terms | What is being said about you that is factually wrong, and which source it traces back to |
Method
How we run it, in enough detail to be checked
The method is published because a measurement nobody can reproduce is not a measurement. You are welcome to run it yourself, and some clients do.
Build the prompt panel
Thirty to sixty prompts drawn from how your buyers actually ask: category shortlisting, problem framing, comparison, price expectation, and direct questions about your company. Every prompt has a written reason for being in the panel, and the panel is then frozen so later runs are comparable.
You get: A written prompt panel with a rationale for each entry
Fix the run protocol
Which assistants, which model versions, which region, how many runs per prompt, and whether memory and personalisation are on or off. Fixing this before the first run is what stops the measurement quietly drifting into a story about whichever conditions produced the nicest chart.
You get: A run protocol another team could follow and reproduce
Run and code every response
Each response is coded for four things: were you named, were you cited with a source, was the description accurate, and which competitors appeared instead. Cited domains are recorded too, because the pattern of who gets cited in your category is often more actionable than your own rate.
You get: A coded response set with the raw transcripts attached
Separate the findings from the noise
We calculate how much variation there is between runs of the same prompt, and treat that as the floor for what counts as a real change. Anything smaller is reported as flat. This is the step that stops a normal fluctuation being sold back to you as an improvement.
You get: A baseline stating sample size and run-to-run variation
Fix what the measurement exposes
Almost every first run surfaces something wrong: a discontinued service still being recommended, an old location, a stale price, a competitor being credited with your work. Those trace back to specific sources, and correcting them at source is the highest-value work on this page.
You get: A source-by-source correction list, with owners
Re-run on a fixed cadence
The same panel, the same protocol, the same coding, quarterly. Model versions are recorded each time so a step change in the numbers can be checked against a step change in the underlying system before anyone claims credit for it.
You get: Quarterly re-measurement against the identical panel
Where we stop
A number without a prompt set, a run count, a model version and a date is not a measurement. It is decoration on an invoice.
We will not produce a single blended AI visibility score, because producing one would require averaging across models and phrasings in a way that destroys the only information worth having.
We will not attribute a movement to our work when the variation between runs of the same prompt is larger than the movement itself. That happens more often than anyone selling this likes to admit, and saying so is the entire reason to trust the times we do claim a result.
And we will not promise a mention or a citation in any assistant, on any timeline, at any budget. The inputs are ours to work on. The output belongs to a system nobody outside it controls.
Questions
What people ask before they trust a number
Can you actually measure whether ChatGPT or Gemini mentions us?
You can measure it as a rate, not as a position. Fix a set of prompts, run each one a set number of times in clean sessions, and record how often your business appears. That gives you something like "named in six of twenty runs of this prompt, on this date, in this model version" — which is a real measurement with a real sample size.
What you cannot do is turn that into a rank. There is no slot to occupy and no leaderboard to climb, and any report that presents one has invented it.
Why do two AI visibility tools give us completely different scores?
Because they are asking different questions. The tools use different prompt sets, different models, different regions, different run counts and different rules about what counts as a mention. Change any one of those and the number moves.
That is not a reason to give up on measurement. It is a reason to insist that whoever reports a number also publishes the prompt set, the run count, the model versions and the date. Without those four things a score is not comparable to anything, including its own previous value.
What is the difference between being mentioned and being cited?
A mention is the model saying your name. A citation is the model attaching a link or a source reference to it. They have very different commercial value and should never be reported as one figure.
A mention with no link builds familiarity and nothing else. A citation puts a route on the page for someone who is already convinced enough to click. In engines that show sources, we track the two separately because the fixes are different: mentions come mostly from corroboration across the open web, citations come mostly from being reachable and quotable right now.
How many runs do you need before a number means anything?
More than most reporting admits to. A single run tells you what happened once, and these systems can return a different set of brands on the next attempt with no change to anything you control.
We run every prompt in the panel multiple times, in fresh sessions with personalisation and memory disabled where the interface allows it, and we state the run count on the report. If a movement is smaller than the variation we see between runs of the same prompt, we say that rather than calling it progress.
Can we see AI traffic in Google Analytics or Search Console?
Partly. Some assistants pass a referrer, so visits from them appear as referral traffic and can be segmented and measured like any other source. That part is real and worth setting up.
A lot of it is invisible. App-based assistants often strip the referrer, and Google reports AI experience impressions and clicks inside its overall Search performance data rather than as a separate breakdown. So referral traffic is a floor on the effect, never the whole of it, and we present it that way.
Can you make a model recommend us?
No. Nobody can, and a guarantee here is a straightforward reason to walk away. Responses vary by user, by phrasing, by region and by model version, and none of that is under any agency control.
What is under your control is being reachable, being described consistently, being corroborated by sources you do not own, and being accurate. Those are the inputs. The output is not something we sell.
Get a baseline with its methodology attached
We will build a small prompt panel for your category, run it properly, and bring you the transcripts along with the numbers. You will be able to see exactly how the figures were produced, which is the point.
Related
Where to go next
- the AI search hubThe wider service this measurement sits inside.
- a one-off AI search auditThe fixed-scope version, if you want the baseline without a retainer.
- what ChatGPT says about your categoryOne assistant, looked at in detail.
- Perplexity and its citationsThe engine where citation and mention are easiest to tell apart.
- analytics and reporting
- the claims we refuse to make
Last updated · Reviewed by Zubair Afzal