How to Measure LLM Visibility: Why Most Brands Track the Wrong Metrics

LLM visibility is a channel you now have to measure, and clicks alone won't capture it. What also counts is whether ChatGPT, Gemini or Copilot name your brand, cite it and recommend it. A brand can shape purchase decisions inside an AI answer long before anyone reaches your site.

AI search is now a channel you have to measure. The way to measure it is to check whether AI answers include, cite, and recommend your brand, which is what LLM visibility measures, not to count the clicks those answers produce. A brand can shape thousands of purchase decisions inside ChatGPT, Gemini, or Copilot without receiving a single visit. Most dashboards still count clicks. The decision happens earlier, inside the response, where your analytics cannot see it.

Here is what you need to know

  • Measure inclusion, citations, and recommendations across a fixed prompt set, and report the result as a share of voice against named competitors.
  • Keep AI visibility and AI traffic on separate dashboards. Combining citation share with referral clicks introduces a measurement error.
  • Diagnose with the four layers: presence, prominence, portrayal, and persuasion. Treat with the four authorities: owned, earned, platform, and entity.
  • Portrayal is the dangerous one. A model that describes you wrongly can cost you more than one that ignores you.
  • Report direction of travel, not decimal scores. AI answers move, and so do the studies measuring them.

Why traditional SEO metrics do not tell the full story

Ranking first in Google does not mean appearing in the AI answer. The two systems draw on different sources, so a page at position one can be absent from the response your buyer actually reads. A scorecard built only on rankings, impressions and clicks is measuring a channel that AI has quietly changed. We have written before about what changes for SEO in the era of AI. Measurement is the part that has lagged furthest behind.

Discovery now happens inside the answer

More buyers get an answer without clicking. SparkToro reports that 68.01 percent of US Google searches in the first four months of 2026 ended without a click. The prospect asks an assistant for a shortlist, reads the options, and forms a view, all before visiting a website. Your brand wins or loses in that moment, and your traffic numbers do not move either way.

Rankings and citations have diverged

The link between rank and citation is weaker than most teams assume. In August 2025, Ahrefs put 15,000 long-tail queries to four assistants and found that only 12% of the URLs they cited also ranked in Google’s top ten for the same query. Tom Capper’s research at Moz, across nearly 40,000 queries, found that 88% of Google AI Mode citations come from URLs that do not rank organically for the same query, a finding later reported by Search Engine Journal. Note that these measure different surfaces. The Ahrefs figure covers standalone assistants, the Moz figure covers Google’s own AI Mode, and the two get quoted interchangeably when they should not be. Different index, different rules, different winners.

You can watch this happen on the term this article is about. We pulled the US results for “llm visibility” in Ahrefs in September 2026. An AI Overview sits above every organic listing and cites a LinkedIn post and two SEO software vendors, none of which hold an organic position on the page. Every page-one organic result is either a tool vendor or a tool roundup. The top organic result receives around 423 visits per month out of 1,200 monthly searches. Position ten gets 36. Regarding a query about measuring AI visibility, the ranked page and the cited page are not the same.

We see the same pattern in client audits. On a recent engagement with a global retail client, the brand held top-three organic positions for its most important commercial terms and was largely absent from AI answers to the equivalent buyer questions. The site was well built. The problem sat elsewhere: the third-party sources the models drew on rarely mentioned the brand.

Search visibility and AI visibility are distinct signals

Treat them as two instruments reading two different things. Tom Capper, Director of Search Product Strategy at STAT and Moz, has put the point plainly: a brand can sit at the top of a traditional results page and still be missing from the AI response its customers are reading. The brands most exposed in our audits are the ones with strong rankings and no view of their AI presence. Rank has become a proxy that the models no longer respect.

What is the primary mistake that brands make when measuring AI visibility?

“One of the biggest mistakes we’re seeing is brands trying to measure AI visibility using traditional search metrics. AI doesn’t always reward the page that ranks first. It rewards the source it trusts most in that moment.” 

Mohammed Faizan, LLM & SEO Lead, Global @ M+C Saatchi Performance

Teams watch AI sessions, AI referrals and AI conversions, conclude the channel is small, and stop. Mentions, citations and recommendations never appear in that view. They still shape the decision.

It is worth being fair about why this happens, because the teams making this mistake are not careless. Clicks were the right metric for twenty years. They are countable, attributable and already wired into every dashboard, board pack and bonus structure in the building. Nobody chose to ignore AI influence. They reached for the instrument already on the desk, and that instrument reads one thing.

The traffic trap

Conductor’s 2026 AEO/GEO Benchmarks Report puts AI referrals at 1.08% of all website traffic, across 13,770 enterprise domains and ten industries. Read that figure alone and you will underrate the channel twice over. The sessions behind it were recorded between May and September 2025, so the number is a floor rather than a current reading. More importantly, the influence happens before the visit, during the questioning stage that qualifies a buyer long before they reach your site.

AI visibility and AI traffic answer different questions

You cannot treat visibility and traffic as the same thing. One tells you whether the model featured you. The other tells you whether someone clicked.

QuestionAI visibilityAI traffic
What it measuresInclusion, citation, recommendationSessions and conversions
InstrumentPrompt panels, share of voice, sentimentGA4, GSC, server logs
When it happensBefore the clickAt and after the click
Best read asBrand and demand signalPerformance signal

Attribution leaks by design

Even when AI does send a visit, the referral chain breaks. Referrers get stripped, links get copied, and much of the traffic arrives as direct. Rand Fishkin’s argument in the same SparkToro analysis is that SEO matters as much as it ever did. It simply no longer earns traffic the way it used to.

Working with clients, we see AI surface first as a rise in branded search and in buyers telling us they came across the brand through AI, well before any of it reaches a session report. Traffic data alone cannot explain AI influence.

What are the four types of AI visibility that matter?

The four are presence, prominence, portrayal, and persuasion. The causal hierarchy comes from the IAB’s Measuring Visibility in the AI Era, published in August 2026, and it gives leadership a shared vocabulary in place of tool-specific scores. Each layer answers one question, then hands the result to the next. We covered the mechanics of how AI assembles those answers in our Performance+ session on GEO and brand discovery.

A word on how to use this. The four layers are a diagnostic. They tell you where you are weak, not what to do about it. Later in this piece we set out the four authorities we build against at M+C Saatchi Performance, which are the levers that move each layer. Diagnose with the layers. Treat with the authorities.

Does the brand appear?

Presence is the floor. If the model never names you, nothing below this line matters. Three metrics measure it, and they are the core of any AI brand visibility program.

  • Mention rate. The share of responses in your prompt set that name the brand.
  • Citation rate. The share of responses that give your own content or domain as a source.
  • Share of voice. Your mentions are divided by all brand mentions across a named competitive set.

How prominent is the brand within the response?

The position within an answer changes its value. Being the first recommendation is not the same as sitting fourth on a shortlist, which is not the same as being mentioned in passing. Users of a shortlist tend to take the first option. Mention is the entry requirement. First position is the prize.

Is the model describing you accurately, and what happens when it is not?

This is the layer most teams skip, and the only one that can actively cost you money. Presence, prominence, and persuasion are all upside problems. Portrayal is a downside.

A model can name you and still damage you. It can place you in the wrong category, quote pricing you retired two years ago, attribute a competitor’s weakness to you, or describe you as the budget option in front of a procurement team building a shortlist. None of this shows up in a traffic report. Most brands find out because a client or a candidate mentions it in passing.

Three things to track per platform:

  • Positioning. Are you the leader, the challenger, the cheap alternative, or the established option?
  • Factual accuracy. Pricing, market coverage, service lines, ownership, headcount.
  • Hallucination rate. How often the model states something about you that is simply untrue.

Run this monthly at minimum. Correcting a misdescription is slow work, because you are not editing a page. You are changing the balance of what the web says about you. The brands that catch it early spend months fixing it. The brands that catch it late spend years.

Does visibility convert into an action?

Persuasion connects visibility to behavior through recommendation strength and downstream engagement. This is where being named starts to influence consideration, preference, and eventually pipeline.

We should be honest about this layer. It is the least reliably measurable of the four today, and anyone selling a clean number for it is ahead of the evidence. Recommendation strength can be scored consistently. The link from that score to revenue still rests on inference rather than attribution. Watch all four layers together, hold this one more loosely than the other three, and expect it to firm up as attribution standards catch up.

What a modern AI visibility dashboard should include

A complete dashboard shows visibility and traffic together rather than replacing one with the other. It should answer, in order, whether you appear, how you are described, how your position compares with named competitors, and only then how much traffic reaches the site. This is the structure we use when auditing brands, and the same structure works for ongoing AI visibility tracking.

Visibility layerWhat to measureCadence and owner
PresenceMention rate, citation rateWeekly, search team
CompetitiveShare of voice against a named setWeekly, search team
AuthorityOff-domain and earned citationsMonthly, PR and content
PortrayalSentiment, framing, hallucination rateWeekly, brand team
CoveragePrompt-set coverage across personasMonthly, strategy
DiscoveryAI impressions from GSCMonthly, analytics
PerformanceAI referral sessions and conversionMonthly, analytics
CommercialBranded-search lift, self-reported attribution, pipelineMonthly, CMO

One rule keeps this readable: traffic sits alongside visibility, never on top of it. Our data and analytics team treats the two as separate reads on separate cadences.

Why earned authority matters more

Earned authority is currently the strongest lever on AI visibility. The models lean far more on third-party validation than on a company’s own pages, so coverage you do not control shapes the answers your buyers read. Muck Rack’s What Is AI Reading? study, May 2026 edition, analyzed more than 25 million links across ChatGPT, Claude and Gemini. Earned media accounted for 84% of citations. Paid and advertorial content accounted for 0.3%.

Third-party sources dominate

Professional platforms, industry publications, community forums, and review sites carry disproportionate weight in AI answers. Semrush, working with LinkedIn, analyzed 325,000 prompts and found LinkedIn cited in 14.3% of ChatGPT Search responses, and in 11% of responses on average across the three platforms it tested. Note the collaboration when you quote that figure. A platform study about the platform is still usable, and the reader deserves to know whose it is. The same trust signal comes from reviews, directories, analyst notes, and expert references. Your press page does not matter. The coverage does.

The four authorities: what to fix when the numbers are bad

The layers tell you that you have a problem. They do not tell you where it lives. We work against four authorities, and a weak layer almost always traces back to one of them.

1. Owned authority

Clear, well-structured pages a model can retrieve, parse and quote without guessing. Weak owned authority shows up as a low citation rate: you are mentioned, but your own site is never the source.

2. Earned authority

Editorial coverage, reviews, analyst notes, and expert mentions from credible third parties. Weak earned authority shows up as a low mention rate. If the sources the model trusts have never written about you, you will not appear in the answer.

3. Platform authority

Presence where your category is already being discussed, from professional networks to the community forums your buyers actually read. Weak platform authority shows up as poor prominence: you get named late, or as an afterthought.

4. Entity authority

A consistent, machine-readable identity across the web, so the model knows who you are, what you do and who you serve. Weak entity authority shows up as portrayal problems: the wrong category, the wrong market, the wrong pricing tier.

Entity authority is the one most often mistaken for a content problem. In a piece on what really drives LLM search visibility, we described an airline client that carried seven variations of its brand name across its digital properties. A model that cannot resolve an entity tends to leave it out. Standardizing the naming and adding organization schema lifted that brand’s AI citation rate by 35% within weeks. No new content was published to achieve it.

Read the two frameworks together, and the diagnosis becomes actionable. A low mention rate is an earned problem. Poor portrayal is an entity problem. Most brands address every AI visibility issue by rewriting their website, which only addresses one of four.

Turn authority into a measurable metric

Track off-domain citation share. It reflects, directly, how much the sources AI trusts trust you. This is the focus of our work on generative engine optimization: the assets other people cite matter more than the ones you publish yourself.

The implication is that you cannot fix AI visibility on your own site alone. Visibility problems usually start in what third-party sources say about your brand. That is why earned authority has become one of the strongest predictors of AI visibility performance, and why any plan to improve brand visibility in AI answer engines has to reach beyond the domain you control.

Directional versus decision-grade

Directional data spots trends and watches competitors. Decision-grade data carries sample size, prompt coverage, reproducibility, and platform range. Only the second kind should move budget. The same test applies to whoever reports it to you, which is why we wrote about how to evaluate a measurement setup before hiring.

DimensionDirectionalDecision-grade
PurposeTrends, monitoring, competitive awarenessBudget allocation, strategy, forecasting
MethodSmall prompt set, single engineLarge prompt set, multi-model, multi-geo
Report toWeekly team standupThe board
Watch forCheap and noisyHigher effort, defensible

Why the numbers move when nothing has changed

AI answers are non-deterministic. Ask the same question twice and you may get two different replies. That is why measuring AI visibility takes discipline rather than a dashboard. Three principles hold:

  • Large, stable prompt libraries beat small prompt samples.
  • Multi-model tracking beats single-model tracking.
  • Precision is not accuracy.

The research moves too, which is the part most people miss. Ahrefs measured the overlap between AI Overview citations and Google’s top ten at 76.1% in July 2025 and 37.9% in March 2026. Some of that gap is a genuine shift in behavior. Some of it is Ahrefs improving its own citation parsing between the two runs, which means the datasets are not directly comparable. If the published studies re-roll on methodology, your internal dashboard certainly will. A monthly move from 18.2% to 18.7% means nothing unless the methodology behind both readings is identical.

Build a measurement framework that leadership can trust

Give leadership a small, defensible set of AI search visibility metrics and KPIs, and bury the vanity numbers. The principles are the ones behind building a unified measurement framework, applied to a surface that did not exist when most reporting stacks were designed.

Report separately: visibility share across a stable prompt set, share of voice against named competitors, sentiment and accuracy, off-domain citation share, branded-search lift, and AI referral conversion.

Cut: single AI ranking positions, raw citation counts to one decimal place, and any blended organic-plus-AI click-through rate.

For teams building this, we recommend continuous measurement across two clearly labeled tiers rather than one definitive score.

What this will not fix

Measurement is a diagnosis, not a treatment. It is worth saying plainly what a good framework leaves untouched.

It will not make a model cite you and tells you that you are not being cited and points to where authority is weak. The work after that is slow, and most of it happens off your domain, where you have influence rather than control.

It will not deliver a clean revenue number this year. The persuasion layer is not there yet.

It will not help a brand with nothing distinctive to say. If the sources a model trusts have no reason to write about you, no amount of prompt tracking creates one.

And it will change. The vocabulary in this piece is roughly a year old. Some of these metrics will look naive in eighteen months, and the teams that fare best will be the ones who wrote their methodology down clearly enough to notice when it stopped working.

“The brands that win in AI search won’t be the ones tracking the most dashboards. They’ll be the ones measuring the signals that reflect real influence and using them to improve visibility where decisions are actually being shaped.”

Mohammed Faizan N, LLM & SEO Lead @ M+C Saatchi Performance

For twenty years, marketers asked one question: how much traffic did we get? AI changes what the answer is worth. The question is no longer whether buyers discovered your brand. It is whether AI recommended it.

Many organizations still treat AI visibility as an extension of SEO. That is the wrong mental model. SEO makes a brand discoverable. AI visibility decides whether it gets recommended. Related problems are no longer the same problem.

So the practical question is not whether to measure AI visibility, but which of your current reports you are willing to stop trusting. Split the two dashboards this quarter. Base visibility on a fixed set of prompts, run it across at least two models and your main markets, and report movement rather than decimals. Then connect that movement to branded search, self-reported attribution, and pipeline so that the figure presented to the board reflects influence rather than activity.

If you would like a second pair of eyes on your setup, talk to us.

Talk to Us

FAQ

LLM visibility is how often, how prominently, and how accurately AI assistants name and cite your brand in their answers. It covers whether ChatGPT, Gemini, Copilot, or Perplexity mention you when a buyer asks a relevant question, whether your own content is given as a source, and how correctly the model describes you. Unlike a search ranking, it is closer to binary: you are in the answer, or you are not. Measure it with a fixed set of buyer questions, repeated on a regular cadence, so you read a pattern rather than a single unreliable reply.

Run a fixed set of high-intent buyer questions across several AI models and your main markets, then record whether the brand is mentioned, cited, and recommended. Turn those records into three figures: mention rate, citation rate, and share of voice against a named set of competitors. Add sentiment and accuracy alongside them. Report movement across weeks and months rather than a single score, because AI answers shift. Keep it separate from traffic reporting, since the two answer different questions.

Yes, and they need separate dashboards. AI visibility is whether the model includes, cites or recommends you, all of which happens inside the answer, before any click. AI traffic is the sessions and conversions that follow. Conductor’s 2026 benchmarks put AI referrals at 1.08% of website traffic, so traffic alone misses most of AI’s effect on the decision. Track both. Never merge citation share with referral clicks into one number, because the merged figure hides almost everything AI is doing to your pipeline.

Because assistants draw on different sources and retrieve them differently. Ahrefs, testing 15,000 long-tail queries in August 2025, found roughly 12% of AI assistant citations also appeared in Google’s top ten for the same query. Tom Capper’s research at Moz, across nearly 40,000 queries, found that 88% of Google AI Mode citations come from URLs that do not rank organically for those queries. Assistants also split a single question into several sub-queries and combine the sources, so a page that ranks moderately across many variations can be cited ahead of the page that ranks first for one variation. Good SEO still helps. It is no longer a reliable indicator of AI presence.

A small, defensible set: visibility share across a stable prompt set, share of voice against named competitors, sentiment and factual accuracy, off-domain citation share, and branded-search lift, with AI referral conversion reported separately. Earned media drives 84% of AI citations, according to Muck Rack (May 2026), making off-domain citation share a priority. Avoid vanity metrics such as a single AI ranking position or raw citation counts to a decimal place, both of which reset every month. Label each metric directional or decision-grade, and let only decision-grade data move the needle.