Why one AI answer proves nothing
Ask an assistant the same question 100 times and there is under a 1 in 100 chance of getting the same list of brands back. Here is what that does to AI visibility measurement, and which numbers survive it.
Someone screenshots ChatGPT recommending a competitor and treats it as a finding. Someone else screenshots it recommending them and puts it on a slide. Both screenshots are close to worthless, and the reason is measurable.
The core finding
SparkToro ran 2,961 prompt runs with 600 volunteers across ChatGPT, Claude and Google's AI answers. There is less than a 1 in 100 chance that an identical prompt run 100 times returns the same list of brands, and less than a 1 in 1,000 chance of the same brands in the same order — even in tightly constrained categories.
Response length varied too, from two or three items to more than ten, with no discernible pattern. The assistants are not returning a ranking that occasionally wobbles. They are sampling from a distribution every single time.
“Any tool that gives a 'ranking position in AI' is full of baloney.”
Rerun the same prompt an hour later and you get a different answer
The most directly relevant study here is Schulte, Bleeker and Kaufmann, who queried ChatGPT, Gemini, Google AI Mode and Perplexity up to ten times each across four campaigns, then compared only the pairs of runs falling within 24 hours of each other so that day-to-day drift could not be blamed.
For identical prompts run within the same 24 hours, the overlap in cited sources averaged 32% to 43%. The overlap in which brands were named ranged from 33% to 48% depending on the category. Under classical search assumptions you would expect near-identical results; instead roughly half to two thirds of the answer changes.
| Engine | Source overlap (Jaccard) | Brand overlap (Jaccard) |
|---|---|---|
| Gemini | 0.505 | 0.409 |
| Google AI Mode | 0.318 | 0.375 |
| Perplexity | 0.282 | 0.492 |
| ChatGPT | 0.233 | 0.437 |
One detail from that study is worth carrying separately: ChatGPT activated web search for only some queries, leaving 57.8% of its runs with no citations at all. A tool checking ChatGPT is often measuring an answer composed without retrieval.
How many runs it actually takes
The same paper answers the practical question with a bootstrap convergence analysis rather than a rule of thumb.
The standard error of a per-brand detection rate falls below 0.10 at 7 runs per prompt, and below 0.08 at 8. Source-level coverage converges more slowly and needs 8. Their recommendation is at least 7 runs per prompt per day for brand visibility, and rolling aggregation over two to four weeks for per-brand estimates that are stable rather than momentary.
They also found prompt-level overlap varying from below 0.2 to above 0.8 within a single campaign. Monitoring on one or two prompts measures the idiosyncrasies of those prompts rather than your visibility, which is an argument for a wide prompt set independent of how many times each is run.
Applied to us as much as anyone: our own free report runs each question once per assistant and aggregates across questions rather than running each prompt seven times. That is enough to answer 'do you appear at all', which is what a business at zero needs to know. It is not enough to estimate a per-prompt detection rate, and we would not present it as one.
Where the noise comes from
A 2026 variance-components study decomposed the non-determinism in brand answers. Its numbers are widely quoted in this category, including by us, and they are worth reporting precisely because they are easy to misread.
| Source of variance | Share of a single response |
|---|---|
| Re-running the same prompt | 34.8% |
| Query language (English vs another language) | 26.5% |
| Brand-in-context interaction | 29.6% |
| Brand × language interaction | 8.6% |
| Brand identity | 1.5% |
| Brand × model and brand × prompt | near zero |
Read that second row carefully, because it is routinely misquoted. It is the natural language the question was asked in — English versus Polish versus Czech — not how the question was phrased. The prompt effect in that study is close to zero. We got this wrong in an earlier version of this guide and it is recorded in the corrections below.
Brand identity — the thing a client pays an agency to change — accounts for about 1.5% of the variance, with an intraclass correlation of 0.0146. Single-answer reliability is near 0.01, rising to about 0.36 across the full crossed design.
That study's measured outcome was multilingual sentiment polarity toward a brand, not whether the brand was recommended or where it placed. The authors describe recommendation presence and ranking as future work. So the 0.01 figure is strong evidence that a single answer is a poor measure of sentiment, and suggestive rather than conclusive about recommendation visibility. The Schulte overlap figures above are the ones that measure recommendation directly, and they point the same way.
What this means for the tools
Set 7 runs per prompt as the floor, and Fishkin's stricter guidance of 60 to 100 runs per prompt as the ceiling for genuine precision, against commercial tools offering 15-prompt or 25-prompt tiers with a single daily run.
Most AI visibility products are sampling well below what their own headline precision implies. That does not make them useless. It makes their decimal places fictional.
Three traps that survive better sampling
Whoever picks the prompts picks the result
The same brand, on the same underlying data, can be shown at 20%, 16.8% or 31.4% share of voice depending on which prompts were chosen and which scoring formula was applied. A provider that selects its own prompt set and then reports improvement against that set has a plain conflict of interest. Ask to see the prompts, and ask whether they changed between reports.
The API is not the product
Answers pulled through an API reflect one account, one geography, one subscription tier and one moment of one model's state. Real users see a distribution shaped by personalisation, location, account state and constant product changes. Anyone measuring this way, ourselves included, is using a proxy and should say so.
Model drift looks exactly like your work
Providers ship model and retrieval updates constantly. One documented case saw GPT-4's accuracy on a fixed task fall from 84% to 51% in three months. A client's visibility can move materially in either direction for reasons entirely unrelated to anything an agency did — and it is genuinely difficult to tell the two apart.
The counter-argument, and what it does not answer
Agencies in this category have a standard response to the reproducibility problem, and it is worth engaging with properly because part of it is correct.
The argument runs: the criticism of tracked prompts is that they are invented questions rather than real conversations, and that objection is weakening. Prompt sets are increasingly built from Google Search Console queries, GA4 landing pages and on-site search terms, and some vendors now license anonymised real prompts from opt-in consumer panels. Scale is not a constraint either — practitioners track hundreds of prompts, not fifteen.
All of that is true, and a prompt set grounded in real buyer language is genuinely better than one somebody made up. But it answers a different objection to the one the research raises.
Better prompt sourcing fixes whether you are asking the right question. It does nothing about the fact that the same question, asked again, returns a different answer. Those are two separate problems, and only the first one has a purchasing solution.
The overlap figures are what make this stubborn. Identical prompts, rerun within 24 hours, agreed on only 33% to 48% of the brands they named. A perfectly sourced prompt still lands in a distribution, and where the prompt came from does not change that.
Scale helps, and less than people assume. Reaching a standard error below 0.10 on a single brand's detection rate takes 7 runs of the same prompt, and stable per-brand estimates take rolling aggregation over two to four weeks. That is a real commitment, and it is a long way from the precision implied by a dashboard reporting share of voice to one decimal place off one daily run.
So the honest position sits between the two. Tracked prompts are not guesswork, and the people saying so are right. They are also not precise, and a report that gives you a number without a sample size and an acknowledgement of variance is overstating what it knows — whatever the prompts were built from.
So which numbers survive?
The line between a defensible measurement and theatre is sharper than the category admits.
| Defensible | Theatre |
|---|---|
| Visibility as a share of runs mentioning you, across many prompts, each run repeatedly, with the sample size stated | Any rank or position |
| Comparative visibility against a fixed competitor set on the same protocol | Share of voice quoted to one decimal place |
| A trend over months, with model drift acknowledged as a confound | Week-on-week movement, 'up two spots' |
| Presence or absence in the consideration set | Attributing a lift to one specific blog post |
| 'This change was undetectable at our sample size' | A single screenshot as evidence of anything |
“Rankings? Probably nonsense. Visibility over many prompts and many runs? Surprisingly useful.”
How this shapes what we do
Two things follow directly, and both are visible in how our own reports are built.
- We spread runs across several questions and several locations rather than repeating one question, because prompt-level overlap varies from below 0.2 to above 0.8 and a narrow prompt set measures the prompts rather than the business.
- We report counts out of a stated denominator and never a score out of 100. A count you can go and check yourself is more use than a composite number we invented, and at these overlap levels a score would be a decimal place laid over a coin flip.
- We say what our sample does not support. Our free report answers whether you appear at all; it does not estimate a per-prompt detection rate, because that needs at least 7 runs per prompt and we do not run that many.
It also sets the honest limit on what any report of this kind is. It is a record of what we observed on the dates stated, not a forecast, and running it again later will produce different figures. That is a property of the systems, not a defect in the method.
Common questions
- Can you get ranked number one in ChatGPT?
- No. There is no stable ranking to occupy. Asking an identical question 100 times returns the same list of brands less than 1% of the time, and the same brands in the same order less than 0.1% of the time. Any product reporting an AI rank is reporting noise.
- How many times should a prompt be run to be meaningful?
- Published guidance suggests 60 to 100 runs per prompt, and adding different wordings and different models buys far more reliability than repeating the same prompt. Most commercial tools sample one to two orders of magnitude below that.
- Is AI visibility measurement worth anything at all?
- Yes, within limits. Visibility measured as a share of runs across many prompts is useful and reasonably stable. Ranks, positions and share of voice quoted to one decimal place are not.
Corrections
What changed in this guide, and why. Listed rather than quietly edited.
- 17 August 2026
A reader checked this guide against the underlying papers and found two errors, both in the variance table, which was the strongest evidence on the page. First, the 32% figure was labelled 'how the question is worded'. It is query language — the natural language the question was asked in — and the prompt effect in that study is close to zero, so the label inverted the finding. The correct figure from the paper's abstract is also 26.5% rather than 32.0%. Second, the claim that wording diversity buys roughly fifteen times more reliability than repetition described language and model diversity, not paraphrasing. Both are corrected above. Separately, we were over-extending that study: its measured outcome is sentiment polarity toward a brand, not whether the brand was recommended, so its 0.01 reliability figure is now presented as suggestive for recommendation visibility rather than as proof of it. The guide now leads on Schulte et al., which measures brand and source overlap directly. The conclusion is unchanged and better supported than it was.
Sources
Every claim above should be checkable. Where a study has limits, they are stated rather than left out.
- Schulte, Bleeker & Kaufmann — 'Don't Measure Once: Measuring Visibility in AI Search (GEO)', arXiv:2604.07585. Four engines, up to 10 repetitions per prompt, 3,409 pairwise source comparisons, 24-hour matched pairs.Grade A/B. Measures brand and source overlap directly. Preprint; methodology and appendices published in full.
- SparkToro (Rand Fishkin) — 600 volunteers, 2,961 prompt runs, 12 prompts across ChatGPT, Claude and Google AI, January 2026Grade A. Large-sample reproducibility study.
- Žatuchin — 'Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers', arXiv:2607.13304, July 2026Grade B for our purposes. The variance decomposition is sound, but its outcome variable is sentiment polarity rather than recommendation or ranking, which the authors describe as future work. Preprint, confidence intervals pending.
- Chen, Zaharia & Zou — documented model drift on a fixed taskGrade B. Illustrates drift as a confound.
More on measuring ai visibility
- Mentions and citations are not the same thing
Being named in an answer and being used as a source are different events with different causes. In our own run, 77 companies were named and 204 domains were cited, and almost nobody managed both.
Want to know where you stand?
We ask five AI assistants for a company like yours and send you a free report showing how often you came up and who came up instead.
Get my free report