
Measure AI visibility with a fixed set of commercially relevant prompts tested repeatedly across selected platforms. Track recommendation rate, relative share, position, citation share, message accuracy and sentiment then connect observable AI visits and self-reported discovery to pipeline or revenue. One-off screenshots and blended scores are not reliable measurement.
AI visibility dashboards often begin with whatever a tool can count. That reverses the logic. A board wants to know whether the brand is entering valuable consideration sets, whether competitors are gaining influence and whether the programme contributes to growth.
The measurement system should answer those questions while preserving enough detail to diagnose movement.
A visibility score is a diagnostic. The commercial question is whether visibility changes qualified demand, pipeline or revenue.
Prompts are longer, more conditional and less standardised than search keywords. Start with real customer language from sales calls, site search, support tickets, reviews and research interviews.
Group prompts by decision stage:
Discovery: Which providers solve this problem?
Fit: Which option suits this customer or constraint?
Comparison: How do named alternatives differ?
Validation: Is the brand credible, secure, sustainable or effective?
Action: Where can the product be bought or how is it implemented?
Assign each prompt a market, persona, category and commercial weight. Keep a stable core set for trend analysis and a smaller experimental set for emerging questions.
The percentage of eligible prompt runs in which the brand appears. Report it by prompt group and platform. A single total can hide commercial weakness.
Formula: brand inclusions divided by eligible prompt runs multiplied by 100.
The brand's inclusions as a share of all tracked competitor inclusions. This shows relative visibility but needs a clearly defined competitive set.
Record whether the brand is first, later in a list or mentioned only in passing. Position is not directly comparable across every answer format so use bands rather than false precision.
Track whether the brand's domain or relevant independent sources are cited. Separate a named mention from a clickable supporting citation. Classify source type and relevance rather than assuming every citation has equal value.
Score a small set of material attributes: category, audience, product facts, geography, price model and differentiators. Flag harmful inaccuracies separately from mild omissions. Sentiment without context can be misleading.
Measure how consistently the brand appears across repeated runs and time periods. A volatile result needs a different response from a stable competitor advantage.
AI influence is partly observable and partly dark. Build an evidence ladder:
Direct: tagged AI referrals and resulting conversions.
Declared: “How did you hear about us?” responses and sales notes naming an AI tool.
Assisted: journeys where an identifiable AI visit occurs before conversion.
Correlated: branded search, direct traffic or pipeline movement aligned with visibility changes.
Experimental: matched markets, held-out prompt groups or staged content releases.
The language should match the evidence. Direct referral revenue can be attributed to that visit under the chosen model. A rise in branded search after a campaign suggests contribution but does not prove causation.
Lead with four lines:
Where did recommendation share change in high-value prompts?
Which competitors gained or lost?
Which sources and messages explain the movement?
What commercial evidence changed at the same time?
Then show actions. A dashboard that reports movement without explaining the next decision is monitoring theatre.
Run tests in a consistent environment and record date, platform, model or product surface where visible, market, session state and prompt text. Fresh sessions reduce personalisation effects but do not eliminate platform variability.
Keep raw answer text or a compliant record for audit. Re-test enough to identify patterns without pretending to know the full distribution of possible answers. When a platform changes materially annotate the series rather than presenting a smooth trend across incomparable conditions.
A composite score can help summarise performance but it obscures the cause. The same score might represent frequent inaccurate mentions or rare high-quality recommendations. Keep the underlying measures visible and document the weighting.
Google now provides a Generative AI performance report in Search Console for its own generative Search features. Use it where available but do not mistake a Google-only view for the whole AI discovery market.
Use enough to represent the real decision journey without diluting the sample with low-value variations. A focused brand may begin with 30 to 50. A multi-market portfolio may need hundreds organised into stable groups.
Weekly or fortnightly suits fast-moving categories. Monthly may be enough for stable B2B markets. Maintain the same cadence when comparing periods.
It usually means the brand's share of mentions or recommendations within a defined prompt and competitor set. Because definitions vary the report must state the exact numerator, denominator and testing method.
Some direct referrals can. Much influence happens without a click or before a later visit so combine attribution with declared and experimental evidence. Do not force an exact ROI from incomplete data.
Google Search Central, generative AI optimisation and measurement guide: https://developers.google.com/search/docs/fundamentals/ai-optimization-guide
OpenAI, ChatGPT search: https://openai.com/index/introducing-chatgpt-search/
Perplexity, platform overview: https://www.perplexity.ai/hub
Lewis et al., retrieval-augmented generation: https://arxiv.org/abs/2005.11401