AI visibility is now measurable, but not with one number

AI-assisted discovery leaves several different kinds of evidence. A page can appear as a cited source, influence a generated summary, receive a referral visit, or help a customer recognise a brand before another channel receives the conversion. These are related outcomes, but they are not interchangeable.

A useful measurement system keeps platform evidence, website behaviour and commercial outcomes separate before interpreting them together. That avoids turning a citation count into a fictional ranking or treating a small referral total as proof that generated answers have no influence.

Start with the platform evidence now available

Google announced dedicated generative AI performance reporting in Search Console in June 2026 and completed its worldwide rollout on 31 August 2026. The reports show impressions for URLs appearing in generative AI features, with views by page, country, device and date. This should now be part of the baseline for any verified website.

Bing Webmaster Tools introduced AI Performance in public preview in February 2026. It reports citation activity across Microsoft Copilot, AI-generated Bing summaries and selected partner experiences, including cited pages and grounding query phrases. Microsoft is clear that the data is sampled, so it is best used for patterns and content decisions rather than an exhaustive total.

OpenAI gives publishers a different signal. ChatGPT search referral links include the utm_source=chatgpt.com parameter, which can be identified in web analytics. Referral visits show successful click-throughs, not every time a brand or page may have appeared in an answer.

  • Google Search Console: generative AI impressions and the pages receiving them
  • Bing Webmaster Tools: sampled citations, cited URLs and grounding queries
  • Web analytics: ChatGPT and other identifiable referral sessions
  • Conversion data: enquiries and valuable actions associated with those visits

Add a controlled prompt sample

Platform reports do not answer every strategic question. A repeatable prompt sample can show how a business is described, which competitors are recommended, what sources are cited and where the available evidence is weak. Build the sample from real customer journeys: problem recognition, approach comparison, provider discovery, due diligence and local intent.

Keep the core prompts, location, account state and collection method stable enough to compare observations over time. Record the date, platform, exact prompt, brand appearance, description, cited sources and any material factual error. Run variants deliberately rather than changing the method by accident.

The result is a sample, not a census. Generated answers can vary between users and repeated runs, and some prompts will trigger web retrieval while others will not. Directional trends become more useful when they agree with platform reports, referral data and improvements in the underlying source coverage.

Measure the information system behind the answer

Visibility metrics explain where a brand appears. Diagnostic metrics explain why it may be absent or misrepresented. Track whether priority pages are indexed, whether important entities and claims are consistent, whether target questions have definitive answers, and whether reputable third-party sources corroborate the organisation's expertise.

For a New Zealand service business, the audit should also test national and local intent separately. A Wellington-based organisation can have strong local evidence while remaining weak for nationwide comparisons, or the reverse. Business Profile accuracy, genuine service-area language, regional references and national expertise all play different roles.

  • Indexation and canonical status of priority pages
  • Coverage of commercial questions and likely follow-ups
  • Accuracy of organisation, person, service and location entities
  • Relevant independent mentions, links, profiles and reviews
  • Consistency between visible content and structured data

Turn reporting into decisions

A monthly scorecard should be short enough to guide action. Report platform visibility, cited or landing pages, qualified referral behaviour, conversions, material prompt observations and the source gaps that will shape the next work cycle. Annotate meaningful releases so that changes can be interpreted against the date they went live.

Prioritise patterns that appear in more than one evidence stream. If a commercially important page earns conventional impressions but never appears in AI citations, examine its passage-level clarity and supporting sources. If citations grow without useful visits or enquiries, test whether the cited content answers an informational need too far from the service decision. Measurement is valuable when it changes the work.