Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.

Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.

Three Sources of Variation, Stacked

Run-to-run variation on the same engine. Disagreement between engines. And churn over time in which domains get cited at all, independent of anything you do. Each is large on its own, and they compound, which is why a single check on a single engine is close to uninformative.

A study of more than 160,000 prompts across four major platforms found roughly a sixth of cited sources shared between them, with only a small percentage cited by all four. Some pairs are closer than others, with search-grounded engines agreeing more than the rest, but nothing approaches the overlap you would need to treat one as a proxy for the others.

Reproducibility within a single engine varies enormously too. Some engines return largely consistent sources across repeated runs of the same prompt; others have been observed replacing the majority of cited domains week to week, and rotating which brand appears first on the same question about half the time.

What This Breaks

Most of the reporting practices that have grown up around AI visibility assume a stability the instrument does not have. Each of these is common and each is unsound given the variation above.

Common practiceWhy it failsWhat to do instead
Checking a prompt once and recording the resultA sample of one from a wide distributionFive runs minimum, report the rate
Reporting a single engine as AI visibilityEngines share only a fraction of cited sourcesReport per engine, never pooled
Comparing this month to last monthBaseline churn swamps the changeSeveral rounds before calling a trend
Screenshotting a good answer as evidenceNot reproducible; the next run may differReport rates, never anecdotes
Blaming a decline on a content changeCannot be distinguished from ordinary churnSay so; see the evidence ladder

The last row is the one with commercial consequences. Selling a content refresh on the strength of a citation decline between two rounds is selling against noise, and it is a pattern worth refusing on both sides of the transaction.

Record the Whole Instrument

An observation without its instrument cannot be compared to anything. Engine, model version, and reasoning mode all change which sources come back, so all three belong in the record alongside the result.

Mode matters more than people expect. The same prompt on the same engine, run in a fast mode versus a deeper reasoning mode, can return largely different cited sources. That means a comparison between two rounds where the mode changed is not measuring your visibility, it is measuring the mode.

Model versions shift underneath you without announcement, which is the other reason to record them. When a number moves sharply, the first question is whether the instrument changed, and a log without version information cannot answer it. That is the same discipline as keeping a change log for site work, argued in SEO split testing.

The Objection That Cannot Be Fully Answered

Assistant answers are personalised by user context and happen in private conversations. No sampled prompt set can be checked against what real buyers actually saw, which means this measurement is structurally weaker than a rank check ever was.

That is the strongest technical argument the sceptics make and it is correct. Your clean-session sample tells you what an assistant returns to a context-free user asking a fixed question. It does not tell you what a specific buyer with history, location, and prior conversation saw, and there is no way to find out.

The reasonable response is not to abandon the measurement but to state the caveat wherever the number appears. A sampled rate with the sample size, the instrument, and an acknowledgement that real answers are personalised is honest. The same number presented as what your customers see is not, and almost nobody in this market ships the caveat.

A Defensible Minimum

If you are going to measure this at all, five runs per prompt per engine, a frozen prompt set, clean sessions, the instrument recorded, and several rounds before calling a trend. Below that, report that the sample is insufficient rather than reporting a number.

That floor is deliberately unglamorous and it is what separates a measurement from a screenshot. It also makes the exercise expensive enough that it should be reserved for decisions that actually depend on it, which for most businesses is a quarterly check rather than a dashboard.

Pair it with the surfaces that are not stochastic. Search Console's generative AI report gives deterministic impression counts, and your own referral data gives arrivals, however undercounted. Building a picture from the stable instruments and using sampling for the part nothing else can see is the sound division, covered in the Search Console AI report and fixing AI traffic attribution.

Questions About Citation Volatility

Why do I get different sources when I run the same prompt twice?

Because these systems are not deterministic and retrieval varies between runs. Research across large prompt sets has found substantial run-to-run change even for identical questions, with some engines far less reproducible than others. A single check is a sample of one from a distribution, not a reading.

Do different AI engines cite the same sources?

Barely. A study of over 160,000 prompts across four major platforms found roughly a sixth of cited sources shared between them, with only a few percent cited by all four. Treating a result from one engine as representative of AI search generally is treating four different instruments as one.

How many runs do I need for a defensible number?

Five per prompt per engine is a reasonable floor for a small operation, and more where answers are unstable. Below that you cannot distinguish a real change from ordinary variation, and the honest output is that the sample is insufficient to estimate a rate rather than a number with false confidence.

My citations dropped between two checks. Should I worry?

Not yet. Cited domains churn heavily at baseline, with large proportions of cited sources changing month to month even without anything happening on your side. A drop between two rounds sits inside that churn. A trend needs several rounds before it means anything.

Does the reasoning mode affect which sources are cited?

It appears to, substantially. The same prompt run in different modes on the same engine can return largely different sources, which means the mode is part of the instrument. Any observation worth keeping should record the engine, the model version, and the mode, or it cannot be compared to anything later.

Primary Sources

SearchHandled Editorial TeamPublished Jan 12, 2026 · Last reviewed Jan 12, 2026. Every factual claim is checked against the linked primary sources; corrections can be submitted through our contact page.