Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.

Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.

The Metric Everyone Reports Is the Wrong One

Appearing in sixty percent of category answers sounds like success. Then you read the answers and most of them say some version of "Brand X exists but lacks the features Competitor Y has". Presence-shaped metrics score that identically to an enthusiastic recommendation.

There are at least four distinct outcomes when your name appears, and the standard visibility number collapses all of them into one. You can be recommended, mentioned neutrally as an option, mentioned with a caveat that undermines you, or explicitly steered away from. Only the first is a win, and the fourth is worse than absence.

The reason the collapse happened is practical rather than conspiratorial. Detecting whether a brand string appears in a response is cheap and scales. Classifying what the response said about it requires reading, and reading is expensive. The measurable thing became the reported thing, which is a pattern worth noticing across this whole category.

Four Outcomes, One Number

Classifying appearances rather than counting them changes what you learn from the same set of prompts. The distribution is the finding, not the total.

OutcomeWhat the answer doesWhat it means for you
RecommendedNames you as the answer, or first among optionsThe win; count these separately
Listed neutrallyIncludes you in a set with no preference expressedConsideration set entry; better than absence, not a result
HedgedNames you with a qualifier that reduces you to a fallbackUsually traceable to a specific source saying that
Recommended againstNames you and directs the buyer to someone elseWorse than absence; investigate the source
AbsentNot mentionedA discovery problem, per the citation gap

The hedged row is the most common and the most actionable, because hedges are specific. "Good for small teams but limited for enterprise" is a claim with a source, and the source is usually a review, a comparison page, or a thread you can find. That makes it a research task rather than a mystery.

Why This Is Worse Than an Ordinary Bad Review

A bad review sits among other reviews and the reader weighs it. An assistant's hedge arrives as a synthesised judgement from a source the buyer treats as neutral, at the exact moment they are choosing, with no counter-argument beside it.

That framing does more damage per instance than the underlying source did. A three-year-old review complaining about a limitation you removed is one voice among many on a review site. The same complaint, absorbed into an assistant's summary and delivered as "users report that it lacks X", is presented as characterisation rather than as one person's experience.

It also compounds silently. Nobody tells you it happened, there is no notification, and your traffic does not necessarily fall, because the buyer never visited to begin with. This is the failure mode that presence metrics are structurally incapable of surfacing.

Measuring It Without Buying Anything

The method is the sampling routine you would run anyway, with one extra column. Capture the full answer rather than a yes or no on whether your name appeared, and classify it.

  1. Reuse your existing prompt set.

    The same buyer questions used for mention-rate tracking, run the same way in clean sessions. The discipline for building and freezing that set is in share of model.

  2. Save the full response, not a flag.

    You cannot classify what you did not keep. This is the single change that separates this measurement from an ordinary visibility check.

  3. Classify into the four outcomes.

    Recommended, listed, hedged, recommended against. Have the same person classify each round, or write the rubric down, because inter-rater drift will otherwise produce a fake trend.

  4. Record the hedge verbatim.

    The exact wording is the lead. Repeated phrasing across engines and runs usually indicates a shared source, which is the thing you can actually act on.

  5. Run enough times to have a rate.

    Single observations are noise given how much these outputs vary between runs, which is quantified in AI citation volatility.

What You Can Actually Change

Not the model. Assistants describe you from what their sources say, so a persistent characterisation traces back to specific third-party content. Correcting that content is the available lever, and it is a slow one.

Start with whether the hedge is true. If an assistant keeps saying you lack something and you do lack it, that is product feedback arriving through an unusual channel, and no amount of source correction changes the answer. If it was true and is not any more, you have a stale-information problem with a findable origin.

Where the source is identifiable, the work is ordinary: get the outdated review or comparison updated, publish something accurate that a system retrieving your category will find, and make sure the correction exists on sources that get cited rather than only on your own site. That is the same mechanism as earning third-party citations, applied defensively.

Questions About Negative AI Mentions

What does being recommended against mean?

Appearing in an answer that steers the buyer elsewhere. The assistant names you, describes you accurately enough, and then explains why a competitor suits the question better. Presence-shaped visibility metrics count that as a mention, which makes it indistinguishable from a win in most reporting.

Why do visibility tools not track this?

Because presence is easy to count and endorsement requires reading the answer. A tool can detect whether a brand string appears in a response cheaply and at scale; classifying whether the response endorsed, hedged, or steered away requires interpreting the text. The easier metric became the standard one.

Is a negative mention worse than no mention?

Frequently yes, because buyers treat assistant answers with something close to the trust they give an expert recommendation. Being absent means you were not considered. Being named with a caveat means you were considered and set aside, in front of someone who had already decided to buy something.

How do I find out how I am being described?

Read the answers rather than counting mentions. Run your category prompts, capture the full response, and classify each appearance as recommended, mentioned neutrally, hedged, or recommended against. It is slower than a dashboard number and it is the only way to see the thing that matters.

Can I get a negative characterisation changed?

Not directly, and not by asking. Assistants describe you from what the sources they retrieve say, so a persistent negative framing usually traces back to specific third-party content: an old review, a comparison page, a forum thread. Correcting the source is the available lever; there is no appeal to the model.

Primary Sources

SearchHandled Editorial TeamPublished Jan 7, 2026 · Last reviewed Jan 7, 2026. Every factual claim is checked against the linked primary sources; corrections can be submitted through our contact page.