Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.
Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.
Start by Expecting Disagreement
Buyers routinely trial three tools, get three different visibility numbers, and conclude the category is unreliable. The disagreement is real and it is mostly correct behaviour: these tools sample a system that returns different results between runs, across engines that cite substantially different sources.
Research across large prompt sets has found major engines sharing only a modest fraction of their cited sources for the same questions. Two tools checking different engines are therefore measuring different things and should disagree. A third checking the same engine at a different time will also disagree, because cited domains churn heavily at baseline.
That reframes the buying decision. You are not looking for the tool with the right number, because there is no single right number. You are looking for the tool whose methodology you can inspect and whose output you could defend, which is a much easier thing to evaluate. The underlying variation is quantified in AI citation volatility.
The Four Questions
Ask these before a demo gets going. They are answerable in one sentence each by a vendor who has thought about measurement, and they reliably separate the field.
| Question | Good answer | What a bad answer tells you |
|---|---|---|
| How many runs per prompt per engine? | A specific number, with sample size shown in the UI | Single runs reported as rates; the number is an anecdote |
| Are sessions clean and unpersonalised? | Yes, and they can explain how | Results contaminated by history, inflated in your favour |
| Do you record model version and reasoning mode? | Yes, on every observation | Trends that cannot survive a model update |
| Do you record how you were described? | Recommended, mentioned, or steered away from | Presence counted as success, per recommended against |
The fourth question is the one that catches most of the category. Detecting whether your brand string appears in a response is cheap; classifying whether the response endorsed you requires reading it. A tool reporting sixty percent visibility that cannot tell you whether those mentions helped is reporting something you cannot act on.
The Pricing Shapes to Watch
Three pricing patterns in this category cost more than the headline suggests, and all three are easier to spot before you buy than after.
Per-engine pricing is the most common. The entry tier covers one engine, and since engines cite different sources, one engine is a partial view sold as a complete one. Work out which engines your buyers actually use, then price the plan that covers those rather than the plan on the front page.
Prompt quotas are the second. A visibility rate depends on repeated runs, so a quota expressed in prompts rather than in responses can mean the tool is checking each prompt once to stay inside it. Ask how the quota is consumed, and whether increasing run counts costs extra, because that is where a defensible methodology gets traded for a cheaper invoice.
Opaque enterprise pricing is the third, and it is a signal rather than merely an inconvenience. In a category where the deliverable is a number whose reliability you cannot inspect, a vendor unwilling to publish what it costs is asking for two acts of trust at once.
Tracking Is Not the Hard Part
The most common complaint about this category is that it produces dashboards rather than decisions. Knowing you are absent from an answer is the easy half. The reason is almost always presence on sources you do not own, which no tracking tool addresses.
That is worth testing during a trial. When the tool shows you as absent, ask what it tells you to do next. If the answer is a generic content recommendation, you have bought a thermometer and will still need to run the diagnosis yourself, using something like the citation gap method.
It also affects how you value the subscription. A tool that measures something you cannot act on is worth its price only if the measurement itself has a use, such as reporting to someone. For many small businesses, quarterly manual sampling produces the same decisions for no monthly fee, which is the honest alternative set out in share of model.
What Nobody Can Sell You
One limitation applies to every tool in the category and almost none of them disclose it: assistant answers are personalised and happen in private conversations, so no sampled prompt set can be checked against what your actual buyers saw.
A clean-session sample tells you what an assistant returns to a context-free user asking a fixed question. It is a reasonable proxy and it is not the thing you want to know. A vendor who states that caveat is being honest about a structural limit; one who presents its number as what your customers see is describing something nobody can observe.
Ranking guarantees deserve the same treatment. No vendor controls whether an assistant names you, so a guarantee of citations is a promise about something outside the seller's control, which is the definition of the warning sign covered in SEO guarantees.
Questions About Choosing a Tool
- Why do AI visibility tools give different answers?
Because they are sampling a non-deterministic system, usually with different prompts, different engines, different run counts, and at different times. Studies of cited sources find major engines sharing only a modest fraction of them, so two tools checking different engines are not measuring the same thing. Disagreement is expected rather than evidence that one is broken.
- What is the single most useful question to ask a vendor?
How many runs per prompt per engine do you average, and do you report the sample size? A tool checking once is reporting a sample of one from a wide distribution. If the answer is vague, or the sample size never appears in the interface, the number is not a measurement and should not be treated as one.
- Should I pay per engine?
Be careful with it. Per-engine pricing is common and it means the headline price covers one surface, with the others as add-ons. Since engines cite substantially different sources, a single-engine subscription tells you about one system rather than about AI visibility, and the real price is the one that covers the engines your buyers actually use.
- Do these tools tell me what to do?
Most do not, and this is the most common complaint about the category. Tracking tells you that you are absent; it rarely tells you why or what to change. Since the reason is usually presence on third-party sources rather than anything on your site, a tool that stops at the number leaves you holding the harder half.
- Is a free AI visibility checker worth using?
As a first look, yes, and it will usually run one prompt once on one engine. That is enough to tell you whether you are catastrophically absent and not enough to establish a rate or a trend. Treat a free scan as a smoke alarm rather than a measurement, and do not let a single result drive a purchase.

