Reader question
How can a B2B company measure AI visibility reliably?
Measure a fixed set of buyer questions across platforms, wordings and repeated runs. Preserve the answer evidence and denominators, then distinguish a visibility observation from a conversion or a causal improvement.
How to Measure AI Visibility Without Cherry-Picking
A practical measurement protocol for mentions, citations, recommendations and repeatability, with worked examples from ARCADE’s CRM study.
Scroll horizontally to inspect the diagram.
“We appeared in ChatGPT” is an observation. It is not yet a measurement system. Before it becomes a performance claim, you need to know the question, the date, the interface, the other attempts and what “appeared” means.
This protocol is intended for a small B2B team that wants evidence it can inspect. It is a proposed operating method informed by research, not a statistical sample-size prescription or a promise of stable rankings.
Short answer
Measure AI visibility with a fixed question set, repeated captures and separate labels for brand mentions, page citations and positive recommendations. Report each numerator with its denominator and preserve missing or failed observations. Segment by buyer need and platform before presenting an overall average.
Write a measurement contract first
Specify the business question: “Does our product enter the shortlist for a founder who needs controlled follow-up?” is different from “Does our educational article receive citations?” Keep the prompt text, target platform, interface or mode, language, relevant location setting, date and session conditions in the record. Use clean conversations where practical and disclose exceptions.
Decide the number of runs in advance. Three runs are a workable exploratory starting point, not proof of representativeness. Decide how to record a clarification, technical failure or incomplete answer. A valid answer with no brand recommendation belongs in the denominator of eligible answers; a failed capture needs a separate availability record.
Keep three visibility measures distinct
| Measure | Positive observation | Not sufficient |
|---|---|---|
| Brand mention | The product is named in answer content. | A similar company name with unresolved identity. |
| Page citation | A visible citation links to the page or domain. | A claim that the platform probably searched it. |
| Recommendation | The answer positively presents the product as an option. | A citation badge or existing-stack example alone. |
Store the exact supporting passage and any ambiguity. A “recommended for this situation” statement is not necessarily an overall winner. Numbered list order is not necessarily preference rank. The citation-versus-recommendation guide works through these edge cases.
Report a denominator a colleague can reconstruct
In a fictional pilot, a product is recommended in six of nine eligible answers. Its observed recommendation presence is 6 ÷ 9, or 66.7%, for that pilot. It is not 66.7% market share or a prediction that two-thirds of buyers will encounter it.
Count a domain once per answer when measuring response-level source presence. Three links to the same domain in one answer are one citing answer, though they may be three citation occurrences in a separate metric. Label both measures rather than letting the larger number stand in for the smaller one.
Use overlap to describe stability, not accuracy
Jaccard overlap is the number of shared items divided by the number of distinct items in either set. In a fictional example, lists containing A, B, C and A, C, D share two items among four distinct items. Their overlap is 0.5. The calculation does not say either list is correct.
For three exact repeats, compare runs 1–2, 1–3 and 2–3, then average within that question. If both citation sets are empty, overlap is undefined: there are no visible sources to compare. If only one is empty, overlap is zero. Do not count two empty sets as perfect source stability.
| Platform | Recommended CRMs | Cited domains | Identical CRM-set pairs |
|---|---|---|---|
| ChatGPT | 0.781 | 0.473 | 42.6% |
| Claude | 0.584 | 0.119 | 7.4% |
| Gemini | 0.646 | 0.171 | 13.0% |
| Perplexity | 0.639 | 0.618 | 11.1% |
Source: ARCADE aggregate findings and methods. Each platform contributes 54 recommendation pairs across 18 questions. Claude’s citation score uses 34 eligible pairs across 13 questions; the other domain scores use all 54 pairs. The study’s selection and annotation limitations apply. These are descriptive platform observations, not quality rankings.
The practical lesson is that source consistency and product-shortlist consistency need separate charts. A similar shortlist is also not the same as an identical shortlist. Do not label a 0.781 mean overlap as a 78.1% probability that an answer repeats exactly.
Subtract the repeat baseline when testing wording
ARCADE compares distance across paraphrases with distance across exact repeats. The current mean difference is +0.0543 on the zero-to-one distance scale. Fifteen of 24 buyer-need/platform blocks are positive, eight negative and one zero. Some positive differences are very small; the sign is not a significance test.
That comparison protects against a common analytical mistake: attributing any changed answer to an edited prompt. Answers can differ without an edit. Your content experiment should likewise record a baseline before changing a page and avoid treating every later difference as evidence of its effect.
Connect reporting to business outcomes carefully
A useful weekly review has three parts: capture availability, visibility observations and business outcomes. Investigate unexpected changes in that order. Missing captures can distort a headline; changed wording can alter comparability; an increase in recommendations may still produce no attributable qualified enquiries.
Google reports its AI-feature traffic within the Search Console Web search type. Do not relabel that aggregate as a clean platform-by-platform GEO report. Keep referral analytics, self-reported discovery and CRM outcomes separately defined.
GEO research demonstrates why visibility needs an explicit metric. Our proposed extension for a commercial team is to preserve a plain ledger rather than depend on a single composite score: question, platform, date, run, capture status, mention, citation, recommendation, evidence, reviewer and notes.
How Gwenth applies this
For Gwenth, a qualified conversation about a relevant sales workflow is more meaningful than an isolated mention in an AI answer. Keep the prospect’s stated source and actual account context separate from assumptions about discovery. The sales execution argument concerns carrying context into action; it does not establish that Gwenth provides the monitoring protocol described here.
Questions about measurement
Can one successful answer prove our content worked?
No. It is useful evidence to retain, but it cannot isolate the cause or show durability. Preserve other runs, unchanged questions and the content change record.
Should we average all platforms?
Only after showing the individual cells and the weighting rule. An overall number can conceal the precise buyer need where the brand appears or disappears.
Method note: ARCADE uses 216 selected answers collected on 11–13 September 2026, including twelve later Claude replacements. Its reused answer pairs are dependent; human semantic validation is incomplete. AI assisted this article’s drafting.
Tags
References
Source material used for factual context in this article.
- ARCADE: September 2026 CRM discovery findings and methods
ARCADE / Gwenth · Accessed
First-party aggregate extract from the 22 September analysis of 216 selected AI answers and 54 Google snapshots. Model-assisted labels; human semantic validation remains incomplete. This is not independent validation of Gwenth.
- GEO: Generative Engine Optimization
Aggarwal et al., KDD 2024 / arXiv · Accessed
Original research introducing a visibility-optimization framework and benchmark with domain-dependent effects. Its experimental results are not a promised uplift in current consumer products.
- AI features and your website
Google Search Central · Accessed
Official Google-only eligibility and reporting guidance; not a ranking recipe for other AI platforms.