Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.
Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.
The Metric, and Why It Is Not a Ranking
Share of model asks a single question: when a buyer describes their problem to an assistant without naming any vendor, how often does your name come back? It is a mention rate across a fixed set of questions, not a position in a list, and treating it as a ranking is the first mistake people make with it.
The distinction matters because there is no stable list to be ranked in. Ask the same assistant the same question five times and you may get five overlapping but different sets of names, in different orders, with different reasoning. There is no tenth blue link to climb toward. What exists is a probability that you appear, and the only honest way to observe a probability is to sample it repeatedly.
That framing also sets the ceiling on what the number can support. Share of model cannot be attributed to revenue, cannot be segmented by buyer, and moves for reasons entirely outside your control, including model updates that reshuffle everything overnight. It answers one question well: are we in the consideration set that assistants draw from? That is worth knowing, and it is all it is worth.
Designing the Prompt Set
Fifteen to twenty-five questions, written the way a buyer would actually ask them, covering the range of intents that precede a purchase in your category. Never include your brand name in a prompt. The moment you do, you are measuring whether the assistant knows you exist, which is a different and much easier question.
Build the set from real language rather than keywords. Sales call recordings, support tickets, and the questions people email you are better sources than a keyword tool, because assistant queries are longer, more conversational, and carry constraints that keyword research strips out. "Best CRM" is a keyword. "We're a five-person agency outgrowing spreadsheets and I need something my team will actually use" is a prompt, and it is the second one that reflects how these tools are used.
| Prompt type | Example shape | What it tests |
|---|---|---|
| Open category | Who should I look at for X? | Whether you are in the default consideration set |
| Constrained | X for a team of five with no technical staff | Whether your positioning survives a qualifier |
| Comparative | Alternatives to the category leader | Whether you register as a credible substitute |
| Problem-first | A description of the symptom, no category named | Whether you are connected to the problem or only the label |
| Local, if relevant | X near a named place, with a constraint | Whether your local presence is legible |
Freeze the set once written. Every prompt you change is a break in the trend line, and the temptation to tweak wording after a disappointing month is exactly how these logs become useless. If a prompt genuinely needs to change, add it as a new one and keep the old one running until you have a year of both.
Running It Without Contaminating the Result
Clean sessions, five runs per prompt per assistant, same week each month, logged immediately. The contamination risks are personalisation, memory, and your own location, and all three inflate your numbers in the direction you would like them to go.
- Use logged-out or fresh sessions.
An account that has discussed your company before will mention your company more. Assistant memory and chat history are the single biggest source of false optimism in DIY measurement, and the person running the test is always the most contaminated account available.
- Run each prompt five times, separately.
New conversation each time, not five messages in one thread. A follow-up in the same thread inherits everything above it, so you would be measuring the conversation rather than the model.
- Log four fields per run.
Was your brand named, which competitors were named, in what order, and which sources were cited if the assistant showed them. The citation field is the most actionable one, because it tells you which third-party pages are doing the work.
- Hold the calendar steady.
Same week, same assistants, same prompts. Model releases land unpredictably and will move your numbers on their own; a consistent schedule at least stops you adding your own variance on top of theirs.
- Record the date and model version.
When a number jumps, the first question is whether the model changed. Without a version note you will spend a week looking for a cause on your own site that was never there.
Reading the Result, Including the Error Bars
With twenty prompts and five runs across three assistants you have three hundred observations, which is enough for a coarse rate and nowhere near enough for a precise one. Report it as a fraction with the sample size attached, and treat month-to-month moves of a few percentage points as noise until proven otherwise.
Expect large differences between assistants, and do not read them as performance differences on your side. Third-party analysis of AI responses has reported that different systems name brands at very different rates overall, with some mentioning brands in nearly every response and others in around half. If one assistant names you far less than another, the likeliest explanation is that it names everyone less, which is why comparing your share against competitors within the same assistant is the only comparison that means anything.
Sentiment is worth logging and rarely worth worrying about. The same body of analysis found the overwhelming majority of brand mentions were neutral in tone, with a small positive minority and a very small negative one. If you find a genuinely negative characterisation repeating across runs, that is a real signal worth chasing to its source. A single unflattering sentence in one run is not.
"Named in 34% of 300 observations across three assistants in July, against 41% in June, on an unchanged prompt set." Not "our AI visibility dropped 17%".
What Moves It, and What Only Appears To
Share of model responds to the same things that make a brand legible to any researcher: consistent identity across the web, third-party corroboration on sources assistants retrieve, and content that answers the constrained question rather than the generic one. It does not respond to files nobody reads.
The citation field in your log is the fastest route to the actual levers. Over a few months it will show you which sites assistants keep reaching for when they answer your category question, and those sites are usually review platforms, directories, and reference pages rather than vendor sites. That matches the broader citation research discussed in comparison and alternatives pages, and it points the work off your own domain more often than agencies selling on-site optimisation would like.
What does not move it: llms.txt, for which no consumption evidence exists, as documented in our review of the file; schema markup added to content that is not otherwise credible; and volume of publishing. What does, in the order we see it hold up, is being indexed at all, being structured so a specific answer is extractable, and being described consistently wherever your name appears. The full sequence is in the AI search visibility playbook.

