Skip to main content

What a Visibility Score Cannot Tell You About How AI Sells Your Brand

Portrait of Anton Sopov

Anton Sopov

Founder, The Prompt Group

Published
Reading time
7 min read
Cover: What a Visibility Score Cannot Tell You About How AI Sells Your Brand

Key Takeaways

  • A visibility score measures whether an AI answer names you. It cannot measure what the answer said about you, and those are different problems with different fixes.
  • Gartner found in May 2026 that only 11 percent of US consumers would let AI make a purchase for them, while 54 percent had to double-check what it told them.
  • Skyword found 47 percent of consumers took significant action on AI-generated brand information: 19 percent avoided a purchase and 17 percent switched brands.
  • When an AI answer contradicts a brand's own claim, 54 percent of consumers go looking for a third-party source and only 29 percent trust the brand.
  • Where your category sits on the Delegation Curve decides which metric matters: an agentic purchase is settled at inclusion, a considered one three questions later.

Most teams tracking AI search are watching one number, and it is the number their tool made easiest to watch. Visibility. What percentage of tracked questions return your name.

It is a real metric. It is also the top of the funnel, and almost every tool in this category stops there, because inclusion is the part that is straightforward to count.

Here is the problem with stopping there. A buyer who sees your name in an AI answer has not bought anything. They have asked one question out of five or six, and the next four decide it. Your score goes up when you are named in the first answer and says nothing at all about the four that follow.

I am going to break down what a visibility score actually measures, what happens in the questions it never sees, and what a complete AI brand strategy has to contain instead.

The Score That Answers the Wrong Question

Every AI visibility tool answers the same question: do I appear in AI. It is a good question. It is not the only one, and for most brands it is not the expensive one.

The second question is how do I appear in AI. Not whether the model named you, but what it said when it did. Whether the price it quoted was right. Whether it described who you are for. Whether it stayed loyal when the buyer named a competitor in the next message.

Those two questions need different prompt sets, different metrics and different benchmarks. Inclusion is measured against competitors. Representation is measured against your own facts.

And most tools blend them: If your branded questions sit in the same set as your unbranded ones, your visibility score is being lifted by questions that already contained your name. A model asked about your company by name will find your company. That tells you nothing about the buyer who has never heard of you, and it makes the dashboard look better than the market does.

If you want to see that split on your own brand, book a walkthrough of the two-question audit.

Inclusion Is Access, Not the Sale

Visibility buys you a place on the shelf. The decision happens later, in the follow-up questions, and the follow-up questions are where the money is.

The research is consistent on this. Gartner's May 2026 consumer survey found that only 11 percent of US consumers would let AI make a purchase decision for them, and that 54 percent of shoppers who used AI had to double-check what it told them. People are not handing over the decision. They are handing over the research, then interrogating the result.

That interrogation is invisible to a visibility score.

There is a useful parallel in medicine. A randomised study published in Nature Medicine in February 2026, with 1,298 participants, found that leading language models identified the relevant condition in 94.9 percent of scenarios when tested alone. When real people used those same models, they identified it in under 34.5 percent of cases, no better than people using ordinary search.

94.9%

of medical scenarios identified correctly by the model alone, against under 34.5 percent when real people used it

Nature Medicine, 2026

The model being right is not the same as the person ending up right. The same gap sits between your brand being named and your brand being chosen.

What the Agent Says About You Once You Are on the List

Once a model has you on the shortlist, four things decide whether you survive the conversation. We call the composite the Agent Promoter Score, and it works the way a Net Promoter Score works, except the promoter is a machine.

Knowledge Accuracy and Factuality: Whether the facts the model states about you are correct. Pricing, hours, coverage, terms, specifications. What does it cost is a question every buyer asks and most brands have never audited.

Subjective Perceived Value: How favourably the model talks about you when the question is an opinion. Is this a good choice for someone like me is not a factual question, and the model will answer it from whatever it can find.

Competitive Loyalty: Whether the model holds its recommendation when a competitor is named in the same sentence. This is where most brands quietly lose, and almost nobody measures it.

Strategic Outcome Effectiveness: Whether the answer moves the buyer toward the next step, with the right links and the right language, or leaves them somewhere that costs you the conversion.

A brand can score well on all four and still have a visibility score of nothing. A brand can have excellent visibility and fail all four. The two numbers are not related, which is exactly why one cannot stand in for the other.

If you want the four pillars scored against your own facts, see how brand optimization works.

Where the Recommendation Actually Dies

Between inclusion and the sale sits the metric almost nobody is watching. We call it Choice Survivorship: how often a brand that gets recommended is still the recommendation once the buyer asks the specific questions.

Price. Fit. Terms. Availability. The unglamorous ones.

Skyword's April 2026 survey of 1,000 US adults found that 47 percent of consumers had taken significant action because of AI-generated information about a brand. Nineteen percent avoided a purchase. Seventeen percent switched brands. This is not a reputational abstraction, it is revenue moving because of a sentence in an answer nobody at the company has read.

The same survey found something worse for anyone relying on their own website to set the record straight. When an AI answer contradicts a brand's own claim, only 29 percent of consumers trust the brand. Twelve percent trust the AI. Fifty-four percent go looking for a third-party source instead.

Your homepage is not the tiebreaker. It is one of three options, and it loses to a stranger.

The Delegation Curve Decides Which Metric Matters

None of this means inclusion is the wrong thing to measure. It means inclusion is the right thing to measure in some categories and badly incomplete in others, and the variable that decides which is how much of the decision the buyer hands over.

That is what the Delegation Curve maps. Every purchase sits somewhere across five degrees of agentic delegation, from a decision made entirely by a person to one made, approved and executed by an agent, and where it sits is set by the stakes of the decision.

The consequence for measurement is direct.

Where the decision sits

What the buyer does

What decides it

What to measure first

Human choice

Never consults a model

Habit, emotion, trust

Nothing in AI yet

AI-assisted and dual choice

Asks five or six questions, decides themselves

The follow-up answers

Agent Promoter Score, Choice Survivorship

AI-delegated

Sets criteria, approves a recommendation

The shortlist and its framing

Inclusion, then representation

Agentic choice

Sets criteria once, the agent buys

Being on the list at the moment of purchase

Inclusion, almost entirely

Which metric leads, by where a category sits on the Delegation Curve.

Read the bottom row carefully, because it is the honest limit of everything above.

In a category where the buyer asks one or two questions and never reconsiders, or where an agent reorders without a human in the loop, inclusion is not an incomplete metric. It is close to the whole game. If you are not on the list when the agent looks, nothing else you did matters, because there is no second question in which to recover. Accenture's 2026 Consumer Pulse research, covering more than 25,000 consumers across 16 countries, found that nearly 60 percent of Canadians would trust a personal AI agent more than their best friend to buy on their behalf, while only 21 percent would let one make the final purchasing decision even inside limits they had set. Both halves of that sentence are true, and the distance between them is where your category is heading.

So the question is not whether to track visibility. It is whether visibility is the only thing you track in a category that has already moved past it.

What an AI Brand Strategy Contains

A visibility dashboard is a measurement layer. A strategy is what you do about the four things it cannot see.

Split the prompt sets: Branded and unbranded questions belong in separate sets with separate baselines. One is measured against competitors, the other against your own facts. Blend them and you will report a number that improves when you do nothing.

Write the knowledge base before the content calendar: Every fact a buyer might ask about, approved and in one place. Pricing, currency, availability, terms, who you serve. Without it, accuracy is an opinion, and the model will form one.

Answer the fit question in public: Publish who your product is for and who it is not for. A model cannot defend you on a question you have never answered anywhere, so it will infer an answer from the safest source it can find. Naming your non-fit is what gives it something to hold onto.

Corroborate externally, then test model by model: Engines repeat what independent sources confirm rather than what your homepage asserts, and each model stocks a different shelf. A brand can survive deep questioning in one and never appear in another.

If you want to see how your brand holds up across all four, start with our AI brand tracking work.

FAQs

How do AI chatbots decide which brands to recommend?
Models assemble an answer from what they can retrieve and what independent sources corroborate, not from what a brand says about itself. That is why external mentions and third-party coverage carry more weight than homepage copy, and why two models asked the same question return different shortlists of different lengths.
Why does my brand appear differently across ChatGPT and Claude?
Each model retrieves from different sources, weighs them differently and returns shortlists of different lengths. A brand can hold a strong position in one engine and be absent from another for the same question, which is why a single blended visibility number across platforms hides more than it shows.
How do I know if AI tools are misrepresenting my brand?
Ask the questions a buyer would ask, then read the answers rather than the score. Start with price, fit, terms and a direct comparison against a named competitor. Most misrepresentation shows up in those four, and none of it appears in a metric that only counts whether you were mentioned.
Is a visibility score still worth tracking?
Yes, and in some categories it is the most important thing you can track. Where a buyer asks one or two questions and never revisits the category, or where an agent purchases without a human in the loop, inclusion effectively settles the decision. The mistake is treating it as sufficient everywhere.
What happens when AI gives wrong information about a brand?
Buyers act on it. Skyword's 2026 research found 47 percent of consumers had taken significant action based on AI-generated brand information, including 19 percent who avoided a purchase and 17 percent who switched brands. Correcting your own site does not reliably fix it, because 54 percent go to a third-party source when an answer and a brand disagree.

Conclusion

Visibility tells you the model knows you exist. It does not tell you what the model says next, and the buyer decides on what comes next.

The brands that will hold their position over the next two years are not the ones with the highest inclusion rate. They are the ones who separated the two questions early, wrote down their own facts before a model guessed at them, and answered the fit question in public while their competitors were still watching a single number go up.

The measurement work is not hard. It is just not the work most teams are currently doing.

If you want to see where your brand sits across both questions and what a broader AI brand strategy would look like for your category, book a call with our team.

Portrait of Anton Sopov

Anton Sopov

Founder, The Prompt Group

Anton Sopov is the founder of The Prompt Group, an AI brand strategy and agentic discovery firm. He works with brands on how they are ranked, cited and described in ChatGPT, Claude, Gemini and Perplexity answers, and is the author of The Delegation Curve.

Both questions, one strategy.

We benchmark where you stand on inclusion and on representation, then build the AI brand strategy that connects both to your business objectives.