22 July 2026

    How to run an AI visibility audit: presence, citations, hallucinations and stability

    Learn how to audit AI visibility beyond mentions: test presence, recommendation role, factual accuracy, citations and response stability across generative search systems. A practical framework for identifying representation gaps and turning them into proportionate action.

    Binocular viewer facing a sunrise, symbolising the monitoring of brand visibility in AI search.
    Matt Noble | Unsplash

    An AI visibility audit does not ask only whether ChatGPT, Gemini, Perplexity or Google AI features mention a brand. It tests whether those systems represent the brand accurately, recommend it in relevant contexts, associate it with the right attributes and distinguish it from competitors.

    This matters because brand visibility in AI is not a single outcome. A company can appear frequently but be described through an outdated offer, placed in the wrong category or confused with another entity. It can also be cited without materially shaping the answer.

    As outlined in our article on brand semantics infrastructure for AI search, a brand needs both a clear entity map and a clear claim map. An audit is where those maps meet the answers users actually receive.

    An audit is not a prompt test

    Asking one model, once, “What are the best providers in this category?” can be a useful observation. It is not a reliable audit.

    Generative systems can vary between runs, prompt formulations, locations, interfaces and models. Google itself notes that AI Overviews and AI Mode may use different models and techniques, so both answers and supporting links can differ. Research on multi-prompt LLM evaluation likewise shows that changing an instruction template can materially affect results.

    The implication is practical: an audit should treat an AI answer as an observation collected under defined conditions, not as a definitive statement about the brand.

    A defensible audit records:

    • the system and interface tested;

    • date, language, market and location;

    • the exact prompt and its scenario;

    • the full response, not only a score;

    • sources or links shown by the system, where available;

    • the criteria used to classify the response.

    This is consistent with the wider distinction explained in GEO after SEO: some elements of AI visibility can be improved directly, some can only be influenced, and some should primarily be monitored.

    What should an AI visibility audit measure?

    The basic unit is not a keyword. It is a scenario in which a person could reasonably ask an AI system for help.

    For a B2B software company, scenarios may include category selection, implementation risk, integrations, compliance requirements, pricing constraints and competitor comparisons. For a hospitality brand, they may involve location, travel purpose, accessibility, amenities, audience type and alternatives.

    Within each scenario, an audit should assess five dimensions.

    Together, these form a representation integrity matrix. It is not a universal industry standard; it is a practical way to prevent a high mention rate from hiding a weak or misleading representation.

    Citation share can be useful, but it should not be treated as proof that a source determined the answer. A displayed citation may support only one detail, while a brand can be mentioned without its own domain being cited. Audit conclusions should therefore combine source data with the response itself.

    Step 1: define the intended representation

    Start by writing down what must be true for the brand to be represented correctly.

    This is not a messaging exercise alone. It is a factual reference model containing:

    • official name and relevant name variants;

    • products, services and categories;

    • audiences, geographies and use cases;

    • claims that are accurate, current and supportable;

    • claims that must not be made;

    • key competitors and common entity-confusion risks.

    For example, “a cybersecurity platform” is too broad if the company only provides managed detection and response for mid-market organisations in selected regions. The more generic the reference model, the harder it becomes to classify an answer as correct or incorrect.

    Step 2: build a scenario panel, not a prompt list

    A useful panel covers the real decision journey rather than every conceivable wording variation.

    It usually includes:

    • informational scenarios: “What is…?” and “How does… work?”;

    • problem scenarios: “How can I solve…?”;

    • comparative scenarios: “Which option is better for…?”;

    • recommendation scenarios: “What should I choose if…?”;

    • reputational scenarios: “Is this company known for…?”;

    • brand-specific scenarios, used to test factual accuracy rather than spontaneous visibility.

    Prompts should be written in the language of the user, not as disguised brand claims. A question such as “Which leading and trusted provider offers…” imports the desired answer into the research design.

    Step 3: separate spontaneous visibility from prompted accuracy

    These are related but different tests.

    A spontaneous-visibility prompt does not name the brand. It asks for recommendations, alternatives or solutions within a defined scenario. It helps answer: does the brand enter the consideration set at all?

    A prompted-accuracy prompt names the brand directly. It asks what the company offers, who it serves, how it differs from alternatives or whether a particular claim is true. It helps answer: what does the system say once the brand is present?

    Combining the two prevents a common error: concluding that a brand has strong AI visibility because the system can describe it after being explicitly given its name.

    Step 4: preserve the evidence

    A dashboard score is not enough to audit a generative answer. Store the full response and the conditions under which it was generated.

    This makes it possible to inspect:

    • whether the brand appeared in the main answer or only in a caveat;

    • which attributes were attached to it;

    • whether competitors dominated the recommendation;

    • whether the answer relied on old, irrelevant or conflicting information;

    • whether a cited source was actually reflected in the wording.

    The audit should also distinguish factual error from uncertainty. “The system did not mention a specific feature” is not the same as “the system falsely claimed the feature exists.”

    Step 5: classify errors by type

    Not every negative result requires the same corrective action. A practical classification includes:

    Step 6: report denominators and uncertainty

    If a brand appeared in six out of ten answers, report 6/10, not only “60% visibility”. If a competitor appeared in nine out of ten, that context matters too.

    Results should be segmented by system, scenario, audience, language and decision stage where the sample allows. A single aggregate score can conceal an important pattern: strong presence in awareness questions but absence from high-intent recommendation scenarios.

    Recent preprints on AI-search measurement argue that citation and visibility metrics should be treated as estimates from variable response distributions, rather than fixed facts. Sielinski’s study and Schulte, Bleeker and Kaufmann’s paper support repeated measurement and cautious interpretation. Both are preprints, so their findings should inform methodology rather than be treated as settled standards.

    Step 7: turn findings into an action plan

    An audit should end with decisions, not just screenshots.

    For each finding, specify:

    1. the evidence;

    2. the likely information gap or risk;

    3. the corrective action;

    4. the owner;

    5. the metric or follow-up test that will assess change.

    Corrective work may involve technical accessibility, clearer product content, claim alignment, internal linking, source corrections, evidence-led editorial content or a stronger external information footprint. It should not default to creating pages for every prompt variation or adding unverified claims to structured data.

    Google’s current guidance is clear: its AI Search features remain rooted in core Search systems, and there are no special requirements or magic optimisations for appearing in AI Overviews or AI Mode. Pages must still be indexed and eligible to show a snippet; strong SEO fundamentals and useful, non-commodity content remain central. See Google’s AI features guidance and generative AI optimisation guide.

    What this does not mean

    An AI visibility audit does not prove how every customer perceives a brand. It does not guarantee future recommendations, rankings, citations, traffic or revenue. It also cannot establish causality every time an answer changes.

    It shows what selected systems generated under documented conditions. That evidence is valuable because it exposes representation gaps that standard rank tracking and analytics may not reveal. But it needs interpretation alongside search data, source analysis and knowledge of the brand’s actual information environment.

    From one-off audit to monitoring

    An audit creates a baseline. Monitoring asks whether that baseline changes over time.

    The measurement core should remain stable: the entity model, fixed scenario panel, classification rules and recording method. Additional exploratory questions can be added around launches, campaigns, rebrands or emerging risks, but they should not replace the comparative panel.

    For recurring research across systems, scenarios and time periods, LLM brand monitoring becomes a useful operational discipline. The goal is not to produce more charts. It is to identify where representation changed, determine whether the change is meaningful and connect it with a proportionate action.

    A well-run AI visibility audit therefore asks a more useful question than “Did the model mention us?”:

    In the situations that matter to our audience, do AI systems represent the brand accurately, relevantly and with evidence strong enough to support the decision?

    If you need a baseline for that question, contact Brand Semantics to discuss the research design.



    Michał Grzebyk
    Michał Grzebyk
    COO Brand Semantics

    Co-founder of Brand Semantics. Engaged in marketing since 2009. Trainer. Strategist. Explorer of new frontiers in modern marketing. Integrates knowledge from diverse fields to deliver innovative business solutions for clients.