Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.
Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.
The Promotion Nobody Notices
A citation rate rises. The report says AI visibility improved. The summary says AI visibility is driving growth. Nothing dishonest happened at any individual step, and the finished claim is two unearned promotions away from the evidence: appearance became attention, and attention became revenue.
This is the characteristic failure of search reporting and it damages the people doing good work more than the people doing bad work. When the revenue does not follow the visibility, the programme loses credibility for a claim it never needed to make, and the genuine results get discounted alongside the overclaim.
The fix is a labelling discipline rather than better analysis. Every number carries the rung it sits on, and a claim may never sit higher than its weakest input. That is cheap to implement and it survives contact with a sceptical client, which is the actual test.
Every number in your reporting can be named as an appearance, an arrival, a demand signal, or an outcome.
The Five Rungs
Each rung is more expensive to obtain and proves more. The useful move is not to climb as high as possible but to know which rung you are on and say so.
| Rung | What it proves | What it does not |
|---|---|---|
| Inspection facts | The page has or lacks something, deterministically | That it affects anything |
| Citations and impressions | You appeared, per the generative AI report | That anyone noticed or acted |
| Clicks and sessions | Someone arrived | That it was worth anything; also undercounted, per privacy measurement |
| Branded search | People sought you out afterwards, per branded search | Which activity caused it |
| Leads and revenue with source preserved | Money, traceable to a first touch | Causation without a control, per split testing |
Note that even the top rung stops short of causation. Revenue attributable to organic first touch is the strongest evidence routinely available and it still does not establish that your work produced it, because nothing was held constant. Causation needs a control group, and most engagements never have one.
Never Bridge Inspection to Outcome
The most common unearned claim in this field takes the form: your entity data is inconsistent, therefore you are under-cited. The first half is an inspection fact. The second is an outcome. The word therefore is doing work nothing supports.
That specific bridge now has a controlled test against it. A study tracking pages that added structured data against matched controls found no meaningful citation lift, which means the mechanism people were reasoning from does not produce the effect they were asserting. The detail is in does schema markup increase AI citations.
The defensible version keeps the two apart: this is hygiene with a sound mechanistic rationale and no measured effect on citations. That sentence is honest, it justifies doing the work, and it does not promise something you would later have to withdraw. Most overclaims in this field can be repaired by that construction.
No Implied Effect Sizes
Nobody has published an effect size for most on-site fixes. "This could improve visibility by fifteen percent" is not a conservative estimate, it is a fabricated one, and it is fabricated in a direction that happens to justify the invoice.
The instruments make it worse. AI visibility measurement has baseline churn large enough to swamp a modest real effect, so even a genuine improvement is frequently unobservable against ordinary variation. Quoting a percentage in that environment implies a precision the measurement cannot deliver, which is quantified in AI citation volatility.
What you can honestly offer is a mechanism and a direction. This fix removes a specific obstacle; sites without that obstacle tend to do better; we cannot tell you by how much and neither can anyone quoting a number. Clients handle that better than vendors expect, because it matches what they already suspect about the field.
Writing a Claim That Survives Scrutiny
Four elements make a claim defensible: what you did, what moved, what instrument measured it, and what you cannot conclude. Most reports include the first two.
- Name the instrument on every number.
Deterministic export or sampled observation. If sampled, the sample size, the engine, and the model version, because without them the number cannot be compared to a later one.
- State the rung the claim rests on.
An impressions rise is an appearance claim. Calling it a growth claim requires the rungs above it, and if you do not have them, the honest report says so.
- Report floors rather than counts where data is partial.
Analytics undercounts by an unknown amount, so "at least" is accurate where a bare number is not, per measuring when visitors decline.
- Separate branded from non-branded.
Branded queries rise for reasons outside the programme, and including them in an acquisition claim attributes somebody else's work to you.
- Write the limits section before the summary.
What you could not measure, what changed in the instrument, and what remains unattributed. Writing it first prevents the summary from quietly promoting anything.
Why the Honest Version Wins Commercially
The instinct is that caveats weaken a report. In practice the opposite holds, because the buyer has been told confident things before and watched them not happen. A report that states its own limits is the first one they have seen that behaves like evidence.
It also protects the relationship at the point of maximum stress. When a number moves badly, a programme that has been labelling its confidence all along can say this sits inside known churn and here is what we do know. A programme that has been claiming precision has to either explain the contradiction or repeat the overclaim in the other direction.
And it is a genuine differentiator in a category where almost nobody does it. The buyers who screen vendors are explicitly looking for people who will admit the field is uncertain and who separate what they inspected from what they sampled, which is a low bar that most of the market fails. The questions they use are in questions to ask an SEO agency.
Questions About Proving Search Work
- What is the evidence ladder?
A ranking of the signals available to you by how much proof each carries. Impressions and citations sit low: they show something happened without showing it mattered. Branded search sits higher. Leads and revenue with the source preserved sit highest. Which rung a number sits on determines what you may honestly claim from it.
- Why does this matter if the numbers all went up?
Because reporting quietly promotes weak signals into strong claims. A citation rate rising becomes 'AI visibility is driving growth', which is two unearned steps: from appearance to attention, and from attention to revenue. When the revenue does not follow, the programme loses credibility it did not need to lose.
- What is the difference between inspection and sampling?
Inspection is a deterministic fact about a page you fetched: the markup is present or absent, the address matches or does not. Sampling is a stochastic observation about what a system returned on a given run. Both are legitimate; stating them in the same voice is not, because one is repeatable and the other has error bars.
- Can I say a fix improved my visibility by a percentage?
Almost never. There is no published effect size for most on-site fixes, and the instruments measuring AI visibility have baseline churn large enough to swamp a modest real effect. Stating an expected percentage implies a measurement nobody has taken, which is the most common overclaim in this field.
- What is the honest thing to say when you cannot prove causation?
Say what you did, what moved, and that you cannot attribute one to the other. That sentence is more credible than a claim the client will later discover was unsupported, and it is the only defensible position when the instrument cannot isolate your change from ordinary variation.

