Measurement

How to measure whether AI assistants are recommending you

There is no Search Console for ChatGPT. A workable measurement system: a prompt set, sampling on a schedule, and the four metrics worth reporting to a client.

Short answer

Measure AI visibility by building a fixed set of prompts your customers would plausibly ask, running each one several times across ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews on a regular schedule, and recording whether you were named, how you were described, which sources were cited, and who appeared instead. Share of answer, the percentage of runs naming you, is the headline number.

The uncomfortable part of this discipline is that the measurement layer barely exists. There is no index status, no impressions report, no position tracking. If you want to know whether the work is landing, you have to build the instrument yourself.

The good news is that the instrument is simple, and running it manually teaches you more than a dashboard would.

Step one: build the prompt set

This is the part that determines whether everything downstream is useful.

Write the questions your customers would type. Not keywords. Questions, in the shape a person uses when talking to an assistant, which is longer and more constrained than a search query.

For a boutique hotel, a starting set might look like:

  • Where should I stay in [city] for a quiet weekend?
  • Best small hotels in [neighbourhood]
  • Boutique hotels in [city] with a pool, under [budget] a night
  • Where to stay in [city] if I want to walk everywhere
  • Hotels in [city] good for a couple's anniversary
  • Is [your hotel name] any good?

Cover four categories deliberately:

Category queries. No brand, high volume, hardest to win. "Best boutique hotels in Fitzroy."

Constrained queries. Where filterable attributes decide it. "Pet friendly hotels in Fitzroy with parking." These are the easiest wins because most competitors have not published the attributes.

Comparison queries. "X versus Y" or "alternatives to X." Worth tracking even when you lose, because the answer tells you who the model considers your peer set.

Brand queries. Your own name. This measures whether the model knows you exist and describes you accurately, which is different from whether it recommends you.

Twenty to thirty prompts is enough for a single-location business. More than fifty and the sampling becomes a chore you stop doing.

Step two: sample properly

Three rules, all of which people break.

Use clean sessions. Log out, or use a fresh browser profile. Turn off memory and personalisation. If you have been researching your own business for weeks, your account is not a neutral instrument.

Set location explicitly. Write the city into the prompt rather than relying on inferred location, unless you are specifically testing local inference.

Run each prompt multiple times. This is the rule most often ignored and the one that matters most. Two runs of an identical prompt can return different businesses. A single run is an anecdote.

Record for each run: date, assistant, prompt, whether you were named, your position in the list, the exact sentence describing you, which sources were cited, and the full list of competitors named.

That last column is the most useful thing in the whole sheet. It tells you who the model considers to be in your category, which is often not who you consider your competitors.

Step three: the four metrics

Share of answer. Across all runs of all prompts, the percentage that name you. One number, trended monthly. Coarse but honest.

Description accuracy. Of the runs that named you, how many described you correctly. A model that recommends you with the wrong price band or the wrong opening hours is actively costing you customers. This metric catches problems no ranking report would.

Citation sources. Which URLs the assistants cite when answering your category prompts. This is a direct instruction list: those are the pages worth being accurately represented on. If a particular city guide is cited in most answers, getting your entry there right is worth more than a month of content.

Competitive set. The businesses named alongside or instead of you, with frequency. Watching this shift is the fastest read on whether your entity work is landing.

Step four: connect it to outcomes

AI visibility is a leading indicator. Tie it to things the business already tracks.

Referral traffic from perplexity.ai, chatgpt.com, copilot.microsoft.com and similar, segmented in your analytics. Real but small, because most people never click.

Branded search volume. Someone reads a recommendation, does not click, and searches your name later. Rising branded search with flat category rankings is a strong signal that recommendations are landing.

Direct traffic and enquiry volume, tracked against the same timeline.

Ask at the point of sale. "How did you hear about us" with an explicit option for an AI assistant. Crude, and currently the only direct attribution that exists.

What a monthly cycle looks like

An hour or two, once a month, in this order:

  1. Run the prompt set, five runs each, clean session, all assistants.
  2. Fill in the sheet.
  3. Update share of answer and description accuracy.
  4. Pull the citation source list and check whether your entry on each cited source is accurate and current.
  5. Note any new competitor appearing repeatedly, and go read why.
  6. Pick the one constrained query you lost most often and fix the missing attribute.

That last step is the whole point. Measurement that does not produce a next action is just a report.

Frequently asked questions

Can I see AI referral traffic in Google Analytics?

Partially. Referrals from Perplexity, ChatGPT and Copilot show up as referral sources when a user clicks a citation link, and you can segment them. What you cannot see is the far larger group who read the answer, never clicked, and later searched your name or walked in. Referral traffic understates AI influence significantly.

How many times should I run each prompt?

At least five, ideally more. Generated answers vary between runs even with identical input, so a single run tells you almost nothing. Five runs gives you a rough hit rate; ten gives you something you can trend.

Should I use my own account or a clean session?

A clean session, with memory and personalisation turned off, and location set explicitly rather than inferred. Otherwise you are measuring what the assistant shows you specifically, not what it shows a stranger.

Are there tools that automate this?

Several commercial AI visibility trackers now exist and more launch every quarter. They save time on sampling and reporting. For a single business you can get most of the value from a spreadsheet and a scheduled hour a month, which also keeps you closer to what the answers say.

Want to know how AI models currently describe your business?

We run a free visibility check across ChatGPT, Perplexity, Claude and Google AI Overviews, then show you exactly which signals are missing.

Book a visibility check

Keep reading