Search results for this topic are full of tools and short on method. That is a problem, because buying a tracker before you have decided what you are tracking produces a dashboard nobody acts on.
This article gives the measurement process first. It works with nothing but a spreadsheet and an hour, and it tells you when a paid tool starts to earn its money.
Why can’t you measure this like rankings?
Rank tracking works because a search result is stable, ordered and reproducible. Ask Google the same question twice and you get broadly the same ten results in the same order.
AI answers are none of those things. There is no position, the sources cited change between runs, and the wording differs every time.
That means the unit of measurement has to change. Instead of “what position are we”, the question becomes “in what proportion of answers do we appear, and what is said about us when we do”.
The variance trap. Asking a model once, seeing your brand, and concluding you have AI visibility is like checking one keyword on one day and declaring an SEO win. Run the prompt five times and you may appear twice.
Metric 1: Citation rate across a fixed prompt set
Citation rate is the proportion of runs across your prompt set in which your brand is mentioned. It is the headline number.
Record it in two forms, because they behave differently. Mention rate counts any reference to your brand. Citation rate counts references that include a link or a named source attribution.
Mention without citation is still valuable, because it puts you in the consideration set. Citation is better, because it sends traffic and is easier to prove.
Score each run as a simple yes or no, then average across runs and prompts. A business starting from zero in a competitive market might reach 10% to 20% within a couple of quarters; the absolute number matters less than the direction.
Metric 2: Share of voice against named competitors
Citation rate on its own does not tell you whether 15% is good. Share of voice does.
For every run, record which brands are named, not just whether you are. Pick five to eight competitors at the start and keep the list fixed.
Share of voice is your mentions divided by all brand mentions across the run set. This is the single most useful number to put in front of a board or a client, because it is relative and it is comparable over time.
It also surfaces surprises. The brand that appears most often in AI answers in your category is frequently not the market leader, it is whoever is best represented in the third-party roundups the models retrieve from.
Metric 3: Sentiment and accuracy
Being mentioned inaccurately is worse than not being mentioned at all, and this is the metric that catches it.
For each mention, record two things. Is the description accurate, and is the framing positive, neutral or negative.
Common failures we see: outdated pricing, services listed that the business stopped offering, wrong location, wrong company size, and confusion with a similarly named business. All of those come from stale third-party sources, which means the fix is off-site rather than on your website.
Log the exact wording of the mention. When you find an inaccuracy you need the sentence to work out where it came from.
Metric 4: Referral traffic from AI platforms
The only metric here that lands in your analytics as a number rather than an observation.
Create a segment or filter for referral traffic from AI platform domains and let it collect. Expect the volume to be small relative to organic search, and expect the quality to be high, because someone arriving from an AI answer has already been told you are relevant.
Two caveats. Some AI traffic arrives without a referrer and gets counted as direct, so treat the number as a floor. And attribution is genuinely hard here, so use it as a trend rather than as a precise count.
How do you build your prompt set?
This is the part everyone skips and the part that determines whether the whole exercise is worth anything.
A good set spans the buying journey rather than clustering on one type of question. Here is a template to adapt, using an AI receptionist business as the worked example.
| Stage | Prompt pattern | Example |
|---|---|---|
| Problem aware | “How do I [solve problem]” | How do I stop missing calls out of hours? |
| Category aware | “What is [category]” | What is an AI receptionist? |
| Comparison | “[Category A] vs [category B]” | AI receptionist vs answering service |
| Shortlist | “Best [category] for [segment]” | Best AI receptionist for a small UK business |
| Local | “Best [category] in [place]” | AI automation agency in the UK |
| Pricing | “How much does [category] cost” | How much does a call answering service cost? |
| Brand direct | “What is [your brand]” | What is 3rive? |
| Brand comparison | “[Your brand] vs [competitor]” | Is 3rive any good? |
Build 15 to 20 prompts on those patterns, covering your main services. Write them in a document, date it, and treat changes as a version change, because a prompt set that drifts cannot show a trend.
How do you run a manual baseline in an afternoon?
- Set up the sheet. One row per prompt per run. Columns: date, platform, prompt, run number, brand mentioned, cited with link, competitors named, sentiment, notes, exact wording.
- Pick your platforms. ChatGPT, Google AI Mode or AI Overviews, and Perplexity cover most of the ground for a UK business. Add Copilot or Gemini if your audience skews that way.
- Use a clean session. Log out or use a private window so that personalisation and chat history do not contaminate the result.
- Run each prompt three to five times per platform. Paste the answer into the sheet, or at minimum record the yes or no scores and the brands named.
- Score consistently. Decide up front whether a mention inside a quoted source counts, and apply the same rule every round.
- Calculate the four metrics. Citation rate, share of voice, sentiment split, and the referral traffic figure from analytics.
- Diarise the repeat. Same prompts, same platforms, same scoring, one month later.
For 20 prompts across three platforms at three runs each, that is 180 observations. Expect two to three hours the first time and about half that afterwards.
What does Search Console show you, and what does it not?
Search Console remains the best free data source here, with real limits worth understanding.
Clicks and impressions from Google’s AI features are folded into the standard performance report rather than presented as a separate source. You cannot filter to “AI Overviews only” and get a clean number.
What you can do is compare query types over time. Segment your queries into informational, transactional, branded and local, then watch impressions, clicks and click-through rate for each group.
The usual pattern is that informational queries hold impressions while losing clicks, because the answer is now given on the results page. That gap is the clearest evidence in your own data that AI answers are affecting you. Google’s guidance in the AI features and your website guide is the reference for what is and is not reported.
When is a paid tool worth it?
Three conditions, and until you hit one of them a spreadsheet is genuinely fine.
Scale. Above roughly 50 prompts, or more than three platforms, manual running becomes the bottleneck.
Frequency. If you need daily or weekly sampling to track a campaign, automation is the only sensible route.
Reporting. If someone other than you needs to see this regularly, a tool that produces the chart saves more time than it costs.
Several vendors sell this, including Semrush, SE Ranking, Profound, Peec AI and Otterly. We use Semrush internally for keyword research, which is a relationship worth disclosing when we assess it. Our full assessment of what these tools do and do not capture is in AI visibility tools compared.
What no tool solves is prompt-set design. The tool automates the running; the thinking is still yours, and a well-run manual baseline beats an automated dashboard built on the wrong twenty questions.