A dashboard reports that a brand has 64% AI visibility. Is that good?
The number may mean that the brand appeared in 64 of 100 responses. It could also be a weighted score combining mentions, citations, position and estimated audience. In another tool, ‘visibility’ may describe how often the brand’s domain occupies a leading citation position. All three calculations can be internally consistent while measuring different phenomena.
This is the central problem with AI visibility reporting: the label often arrives before the operational definition.
AI visibility is not one property of a brand. It is a family of observations about whether a brand appears, the role it receives, whether it is recommended, which sources are shown, whether claims are accurate and how stable the result is. A composite score can summarise those observations, but it cannot replace them.
The practical question is therefore not ‘What is our AI visibility score?’ It is:
What event did we measure, across which research panel, using what denominator, and which decision can the result support?
The same label can hide different measurements
There is no single, provider-defined standard for ‘AI visibility’. Public tool documentation illustrates the problem.
Semrush’s AI Visibility metrics use ‘AI Visibility’ for a 0–100 competitive benchmark based on how often a brand appears in generated answers. On the same documentation page, ‘Visibility’ in Prompt Tracking refers to a domain’s progress in top citation positions for tracked prompts. OtterlyAI’s Brand Visibility Index combines brand coverage with a measure derived from average position.
These methods are not interchangeable, although each can answer a useful question. The error begins when their outputs are compared as if they measured the same event, population and probability.
Start with the unit of observation
Before choosing a formula, define what constitutes one observation:
– a run is one execution under specified conditions;
– a response is its output;
– a prompt is the submitted wording;
– a scenario is the underlying information or decision need;
– a citation is one visible reference;
– a claim is one verifiable statement.
Ten mentions in one response are not equivalent to presence across ten scenarios. Three links from one answer do not create coverage across three answers. A prompt is wording, while a scenario represents an intent, audience and context.
A reproducible AI visibility audit therefore starts with entities, intent scenarios and controlled research conditions.
Presence, mentions and coverage are not always the same
Presence is the simplest dimension: did the brand appear in the response?
A response presence rate can be defined as:
valid responses containing the brand / all valid responses in scope
This is binary: a response counts once whether the brand appears once or twelve times. A mention count may instead record every occurrence. Mention share can divide one brand’s occurrences by all occurrences of the audited competitor set, but is sensitive to response length, repetition and extraction rules.
Scenario coverage answers a different question:
scenarios in which the brand met a defined presence threshold / all eligible scenarios
The threshold matters. If ‘mention rate’ and ‘presence rate’ use the same binary calculation, they are simply two names for one metric.
The 5P AI representation audit model places these measures under Presence, but deliberately keeps Position, Provenance, Precision and Persistence separate. Presence establishes inclusion. It does not establish meaning.
A mention does not identify the brand’s role
The same brand can appear as a primary recommendation, shortlist option, category example, information source, comparison point, unsuitable option or confused entity. All increase binary presence. Only some indicate genuine consideration.
Representation therefore needs classification, not only sentiment. A positive sentence can place a brand in the wrong category, while a neutral answer can identify it accurately as the best fit for a constrained use case.
Recommendation needs its own denominator
Recommendation rate should not normally be calculated across every prompt in a study. Definitional, troubleshooting, comparison and explicitly branded questions do not all create a genuine opportunity for recommendation.
A more interpretable formula is:
valid eligible responses recommending the brand / all valid eligible responses
Eligibility must be defined before viewing results. The numerator also needs a rule: does shortlist inclusion count, and is the first recommendation weighted more heavily than the fifth? A B2B procurement study may treat shortlist inclusion as meaningful, while a consumer study may require explicit endorsement. The metric becomes credible when the rule is declared and applied consistently.
Citation metrics answer a different question
Citation rate measures visible source inclusion, not brand recommendation.
At response level, it may be defined as:
valid responses citing the audited domain / valid responses capable of displaying citations
Citation share asks what proportion belongs to a domain or source class. Citation prevalence asks how broadly a domain appears across responses. Platforms differ in citation volume, presentation and use of web search.
A visible citation establishes visible provenance. It does not prove that the page shaped every claim. Assessing support or contribution requires claim-level analysis:
Count citations as citations. Do not silently convert them into recommendations, influence or factual support.
Representation cannot be reduced to favourable wording
Brand representation covers the category, offer, audiences, capabilities, limitations and relationships expressed by the answer. Measuring it requires a verified reference layer. A claim map can identify statements that are correct, outdated, false, unverifiable or missing a material limitation. Without it, ‘hallucination rate’ is often only an impression.
A claim accuracy rate can be expressed as:
correct atomic brand claims / all atomic brand claims eligible for verification
An incorrect office location and an invented capability both count as errors, but their consequences differ. Preserve error type and severity alongside the rate.
Brand semantics infrastructure connects measurement to entity maps, claim maps and source alignment, defining what a correct representation would contain and how it can be verified.
The denominator can change the conclusion
Every percentage contains a policy decision about what belongs under the line.
Suppose 100 runs were scheduled. Five returned provider errors, five were empty and ten did not activate a citation-capable search mode. If the brand was cited in 18 answers, its reported citation rate could be:
– 18/100, if every scheduled run remains in the denominator;
– 18/90, if technical failures are excluded;
– 18/80, if the metric is limited to valid, citation-capable responses.
Each figure describes a different process. Technical errors are not neutral answers or brand absence, but silently removing them can make an unreliable platform look cleaner. Report completeness separately and declare every denominator.
At minimum, publish both the percentage and the fraction: 18/80, not only 22.5%.
Platform averages are not neutral
Combining ChatGPT, Google AI Mode, Gemini, Claude, Copilot and Perplexity creates choices about platform weights, equivalent product surfaces, search modes and classification rules.
An unweighted average gives the mean across selected platforms. It does not estimate the probability that a customer will encounter the brand, which would require credible audience and scenario weights.
Aggregation can also reverse an apparent trend. A brand may improve within every stable segment yet decline overall if the second measurement contains harder scenarios or gives more weight to a low-performing platform.
Platform-level and scenario-level results should therefore remain available beneath any total.
A percentage is an estimate, not a fixed platform property
Generative answers vary across runs, wording and time. Two recent preprints provide evidence, although neither is a universal benchmark for every platform or market.
Quantifying Uncertainty in AI Visibility studies repeated citation measurements across Perplexity Search, OpenAI SearchGPT and Google Gemini. It frames citation metrics as sample estimates of an underlying response distribution and finds that some apparent differences between domains fall within measurement noise.
Don’t Measure Once analyses repeated brand and source observations across several AI search engines. It likewise argues that visibility should be treated as a distribution rather than a single-point outcome.
Both are 2026 preprints rather than established cross-platform standards. Sielinski is affiliated with commercial AI visibility company IQRush, while Schulte discloses an affiliation with Aurora Intelligence alongside his university role. Their methods and limitations matter more here than their headline figures.
Sample size depends on the metric, scenario, platform and variability. One run per prompt cannot establish persistence. Material decisions should include the fraction, number of runs, a stability measure and, where justified, an uncertainty interval.
When a composite score is useful
A composite score can provide an executive signal, track direction under a stable method and trigger investigation. The problem is not aggregation itself, but aggregation without traceability.
The OECD and European Commission handbook treats composite-indicator construction as a methodological task. Selection, normalisation and weighting can change the output.
For AI visibility, components need operational definitions, weights must be explicit, panel changes and technical exclusions must be documented, and component results must remain accessible. The score must not be presented as market share, user exposure or revenue impact.
A score should be a navigation layer, not a barrier between the reader and the evidence.
The AI visibility metric contract
Brand Semantics proposes an AI visibility metric contract: a compact specification that should accompany every material metric.
Contract element | Required question |
|---|---|
Decision | What decision is this metric intended to support? |
Event | What exactly is counted: presence, mention, recommendation, citation, role, claim or error? |
Observation unit | Is one observation a run, response, prompt, scenario, citation or claim? |
Numerator and denominator | What sits above and below the line? |
Eligibility and exclusions | How are technical errors, empty responses and ineligible scenarios handled? |
Research panel | Which platforms, models, modes, markets, languages, audiences and scenarios are included? |
Aggregation | Are segments weighted, and how? |
Uncertainty and comparability | How many runs were collected, how variable were they and did the method remain stable over time? |
This is a Brand Semantics methodological proposal, not an official platform standard. It makes a metric auditable before it becomes persuasive.
For example, a recommendation-rate contract could specify explicit positive recommendation at response level, 40 predefined non-branded purchase scenarios, five runs per scenario, equal scenario weights and separate platform results. Provider failures would be excluded from the rate but reported separately. Periods would be compared only if the scenario library, model surface and classification rules remained stable.
Report a scorecard, not only a score
A useful measurement system has three connected layers.
Management view – a small set of directional indicators, material risks and changes over time.
Diagnostic view – results by platform, market, audience, scenario type, recommendation role, source class and error category.
Evidence view – full responses, citations, claims, classifications, timestamps and technical statuses.
This complements the process for interpreting and reporting an AI visibility audit. It also reflects the GEO control surface: brands can optimise controlled assets and influence information conditions, but mentions, citations and recommendations remain observed outcomes.
Semantio supports this approach by preserving scenario-based, cross-model monitoring and separating answers, recommendations, sources and errors. A tool does not remove the need for metric definitions; it can help apply them consistently and retain the underlying evidence.
What this does not mean
One synthetic score is not forbidden. It can be useful when its method is stable, transparent and decomposable.
Every project does not need every metric. A source audit, recommendation study and factual-risk review require different scorecards.
Citations and recommendations are not automatically valuable or correct. A citation may be outdated, while a recommendation may rely on a false capability.
More observations do not fix poor research design. Repeating an unrepresentative prompt panel produces a more precise estimate of the wrong construct.
Tools are not directly comparable by default. The same label may hide different events, panels and weights.
AI visibility is not market share or business impact. Connecting representation to demand, conversion or revenue requires additional evidence.
Measure the event before naming the score
AI visibility becomes useful when it is decomposed into answerable questions: does the brand appear, in which scenarios, is it recommended, what role and category does it receive, which sources are visible, are its claims correct and does the result persist?
A single number may summarise part of this picture. It should not be allowed to redefine the picture.
Before accepting an AI visibility score, ask for its contract: event, unit, denominator, panel, exclusions, aggregation and uncertainty. Without them, the number may decorate a dashboard, but it is not strong enough to guide a decision.
Brand Semantics applies this distinction through AI Strategic Consulting, connecting measurement design with entity, claim, source and representation analysis.
