Why every AI visibility tool gives you a different number
Why every AI visibility tool gives you a different number for the same brand, and the sampling method that produces a figure you can defend.
Short answer
Every AI visibility tool gives you a different number because the thing it measures is unstable. Language models answer differently each time they are asked, most dashboards ask once, each vendor runs its own prompt list, and ChatGPT, Perplexity and Google AI Overviews retrieve from different indexes. Blending that into one score hides the variance instead of reporting it.
Ask three tools how visible your brand is in AI answers and you will get three numbers. That happens because the models answer differently on every run, each tool samples a different number of runs against a different prompt list, the engines retrieve from different indexes, and the result is then folded into one score. There is no single number to be right about. There is a distribution, and almost nobody reports it.
Why every AI visibility tool gives you a different number
The disagreement comes from four places, and every tool on the market handles at least one of them badly.
The answer changes every time you ask
Large language models are non-deterministic. Send the same prompt to ChatGPT twice and you can get two different lists of businesses. This is how the models generate text, not a bug a vendor can patch. A single run of a prompt is one draw from a distribution, and any tool that reports one draw as a score is reporting a coin flip with a decimal point.
Published sampling work on LLM non-determinism puts numbers on this. Five repetitions of a prompt carry a margin of error around plus or minus 27 points. At that width, most month-on-month movement is indistinguishable from noise. Twelve or more repetitions per prompt is the floor at which the figure starts to hold still.
Most dashboards ask once
Repetitions cost money. Every extra run is an API call, so the commercial pressure on a visibility tool is to run each prompt once, store the result, and draw a trend line through the noise. Check the methodology page for the repetition count. If it is not there, assume once.
The prompts are not your prompts
Each tool ships its own prompt library. One vendor asks "best specialty coffee in Lisbon". Another asks "where should I get coffee in Lisbon". A third generates variants automatically and does not show them to you. Different questions get different answers, and the retrieval behind them differs too. Ahrefs found roughly a third of AI Overview citations come from pages ranking outside Google's top 100 for the parent query, because the engine fans out into sub-questions you never see. Two tools with two prompt lists are sampling two different sets of sub-questions.
The surfaces are not the same product
This is the one that does the most damage. Google's AI Overviews and AI Mode retrieve from Google's index with Googlebot, which renders JavaScript. ChatGPT, Perplexity and Claude fetch with their own crawlers, and in the Vercel and MERJ analysis of more than 500 million fetches, none of those crawlers executed JavaScript. A site that builds its content in the browser can be fully visible in AI Overviews and invisible to ChatGPT on the same day.
Local results diverge further. Ben Fisher's August 2026 test of 2,880 prompts found ChatGPT grounding on Yelp in roughly 96 percent of local runs, after OpenAI's licensing deal with Yelp. Google's AI surfaces retrieve from Google's own index, which we covered in Google AI Overviews for local businesses. A cafe with a strong Yelp profile and a thin presence in Google's index will score differently on each surface.
A tool that averages those surfaces into one figure is averaging a Google measurement with a Yelp measurement and calling it AI visibility.
Blended scores hide all of the above
Put the four problems together and you get the typical dashboard. A handful of vendor-chosen prompts, run once each, across several engines with different retrieval, weighted by an undisclosed formula, presented as a percentage with a trend arrow. The arrow moves because the noise moves. The number is not wrong so much as unmeasured.
Citations and mentions are not the same count
Even a careful tool has to decide what it counts, and the two candidates diverge sharply. Semrush's ghost citations study found 61.7 percent of citations produced no brand mention at all. ChatGPT cited a source in 87 percent of answers and named a brand in 20.7 percent.
So a tool counting citations will report a far higher figure than one counting mentions, for the same brand on the same day. If you sell something, the mention is the outcome you want. A link in a footnote is a different thing, and Pew Research found only 1 percent of users clicked a link inside an AI summary. Ask any tool which of the two it counts. If it cannot answer, that is your answer.
How to measure AI visibility honestly
We do this for clients and none of it needs a subscription. It needs discipline.
Fix the prompt panel and never let it drift
Write 20 to 40 prompts a real customer would type, in the phrasings they use. Freeze the list. Add prompts only at a dated boundary and report the old and new panels separately. A panel that changes every month cannot show a trend, because you changed the ruler.
Run each prompt at least 12 times
Per surface. Same day, spread across hours, because time of day is another source of variance. Record every raw answer, not the tally, so you can re-score later.
Report the spread, not the point
Publish the range and the count. A line like "named in seven of twelve runs on this prompt, and in between three and ten across the panel" is a measurement. A single percentage is a summary that hides whether it was stable or a lucky week. When a change lands inside last month's range, say so and say nothing else.
Keep Google surfaces and ChatGPT in separate columns
Never blend. Google AI Overviews, AI Mode, Copilot, ChatGPT and Perplexity are different retrieval systems and they should be different columns. This is the split we use.
| Surface | Retrieval source | First-party report |
|---|---|---|
| Google AI Overviews and AI Mode | Googlebot, renders JavaScript | Search Console Generative AI report, impressions only |
| Microsoft Copilot | Bingbot, renders JavaScript | Bing Webmaster Tools AI Performance report, citations and grounding queries |
| ChatGPT search | OAI-SearchBot, no JavaScript, Yelp for local | None |
| Perplexity | PerplexityBot, no JavaScript | None |
| Claude | Claude-SearchBot, no JavaScript | None |
Use the first-party reports where they exist
Two of those rows have real data behind them and most tools ignore both. Bing Webmaster Tools launched an AI Performance report in February 2026 that shows Copilot citations and the grounding queries the model generated, which is the only place you can read the sub-questions directly. Google Search Console's Generative AI report went global in August 2026 and shows impressions only: thin, but Google's own count rather than a sample. Both are walked through in how to track AI search visibility.
Separate mentions from citations from traffic
Three counts, three columns. Mentions are what you want. Citations are what tools usually count. Traffic is what the finance team asks about, and it is small: Conductor's benchmark puts AI referrals around 1 percent of traffic, and Ahrefs reported 0.5 percent of its own traffic from AI alongside 12.1 percent of signups. Attribution is weak because many AI referrals arrive with no referrer and land as Direct, so the analytics figure is a floor rather than a count.
What to do with the tool you already pay for
Keep it if you like the interface, but read its number as one sample. Ask the vendor three questions in writing: how many repetitions per prompt, whether the score counts mentions or citations, and whether Google surfaces are blended with ChatGPT. A methodology page that omits the repetition count is telling you the count is one.
The mechanics of getting named in the first place are covered in how AI assistants choose which businesses to recommend. Measurement has to come first, though. A tactic tested against a single-run dashboard will look like it worked about half the time regardless of what it did.
Twelve runs, a frozen panel, separate columns, a range instead of a point. That produces a number you can defend in a meeting, and a number you can defend is the only kind worth paying for.
Frequently asked questions
Which AI visibility tool is the most accurate?
None of them can be, in the sense the question implies, because the underlying answer is not fixed. A tool that runs each prompt once reports noise. Judge a tool on whether it publishes its prompt list, its repetition count and a range, and whether it keeps Google surfaces separate from ChatGPT. A single blended score with no variance is a chart, not a measurement.
How many times should I run each prompt?
At least 12 repetitions per prompt, per surface. Published sampling work on LLM non-determinism puts five repetitions at roughly plus or minus 27 points of margin, which is wide enough to turn a real gain into a loss on the next run. Twelve or more narrows that to something you can track month to month. Fewer runs are cheaper and mostly meaningless.
Can I get this data from Google or Microsoft directly?
Partly. Bing Webmaster Tools has an AI Performance report, launched in February 2026, that shows Copilot citations and the grounding queries the model generated. Google Search Console added a Generative AI report globally in August 2026, and it shows impressions only. There is no equivalent for ChatGPT, Perplexity or Claude, so those surfaces still need prompt sampling.
Does being cited mean the AI recommended my business?
No. Semrush's ghost citations study found 61.7 percent of citations produced no brand mention. ChatGPT cited a source in 87 percent of answers and named a brand in 20.7 percent. A tool that counts citations and calls them recommendations is inflating the figure. Track mentions and citations as two separate columns and report both.
Want to know how AI models currently describe your business?
We run a free visibility check across ChatGPT, Perplexity, Claude and Google AI Overviews, then show you exactly which signals are missing.
Book a visibility check