Article / A score is not a measurement plan

How to Measure AI Search Visibility Without Making Up a Score

There is no honest universal AI rank. Measure citations, repeated prompts, referrals, conversions, and revenue separately, with every limit visible.

The market loves one big number because one number is easy to sell. The problem is that AI citations, observed prompts, referral sessions, assisted conversions, and revenue come from different systems with different blind spots. I would rather show the relevant measures separately than collapse them into a polished score that nobody can audit.

The measurement rules that keep the report honest

  • Choose the business decision before choosing the metric.
  • Report recurrence and ranges across repeated samples, not one lucky screenshot.
  • Never turn incompatible platform reports into a cross-platform score without exposing the formula and limits.

Each metric answers a different question

A useful report starts by deciding which question matters. It then uses the narrowest record that can actually answer it.

Each metric answers a different question
MetricQuestion it can answerImportant blind spot
Official citation dataWhich pages a supported platform reported citing, how often, and for which sampled grounding queries when available.It may not show placement, every answer, authority, traffic, or other platforms.
Controlled prompt observationsHow often the site appeared for a frozen prompt panel under recorded test conditions.It measures the panel, not the complete market or every personalized answer.
Referral sessionsWhich tagged or recognizable AI sources sent visits, where they landed, and what they did next.Zero-click exposure, stripped referral data, and the exact citation that caused the visit.
Conversions and revenueWhat qualified actions and completed business outcomes were recorded after those sessions.Incremental lift or causation without a stronger experimental design.
A custom indexMovement inside a disclosed, fixed internal model when the same inputs and weights are used over time.Objective market share or comparability to another vendor with different prompts, weights, and coverage.

The first problem is the denominator

To calculate true visibility, we would need to know every relevant AI answer in which a site could have appeared. No official source gives us that cross-platform denominator. A tool chooses a prompt set, chooses the platforms, chooses the markets, chooses the repeat count, and then decides how to weight the results.

That can still produce a useful private index. It becomes misleading when the choices disappear and the output is presented as an objective percentage of the whole AI search market. I want the prompt universe, sample size, platforms, locations, dates, and formula visible before I accept the score.

  • Publish the exact prompt categories and sample size.
  • Name the supported products, markets, languages, and collection dates.
  • Expose every weight used to combine unlike measurements.

Use official reports for what they actually expose

Google documents AI Overview and AI Mode activity inside the overall Search Console Web search type. It does not document a separate AI-only performance segment in that report. That means a Google Web click or impression cannot simply be relabeled as AI traffic.

Bing Webmaster Tools has an AI Performance preview with citation totals, average cited pages, citation trends, cited pages, and sampled grounding queries across supported Microsoft surfaces and integrations. Bing also states clear limits: a citation total does not show placement, average cited pages is not authority, and grounding queries are a sample.

  • Keep Google Web performance labeled as combined Search activity.
  • Keep Bing citation data labeled as Microsoft ecosystem data.
  • Carry the platform's stated caveats into the client report.

A prompt test needs a frozen sample

Prompt observation is valuable when the official report cannot answer the question. It is also easy to manipulate. A tester can add favorable prompts, remove failed runs, switch locations, change account state, or stop after one good screenshot.

I freeze the prompt panel before collection, run every prompt the same number of times, and save every result. I report observed citation rate as citations divided by completed runs, with the sample size beside it. I do not call that total market share.

  • Separate definitions, comparisons, recommendations, troubleshooting, and current-fact prompts.
  • Record platform, surface, locale, device, account state, and exact timestamp.
  • Report missing answers, failed runs, brand errors, and negative outcomes.

Referral data is a different record

OpenAI says ChatGPT adds a chatgpt.com UTM source to referral URLs, which can make attributable sessions visible in analytics. Other products and journeys may pass different information or no useful referral at all. A referral report therefore measures identifiable visits, not every AI-assisted discovery.

In GA4, I use session-scoped traffic acquisition data for the visit and keep it separate from first-user acquisition. I then break it down by landing page, engagement, key event, and revenue. If the source was inferred through a custom channel rule, I label the rule and do not mix it with an explicit platform tag.

  • Preserve the raw source and medium before grouping channels.
  • Keep session acquisition separate from first-user acquisition.
  • Report landing pages and qualified actions, not sessions alone.

The business outcome belongs in its own column

A citation can matter without generating a click, and a click can matter without producing revenue on the first session. I still refuse to collapse those stages. The report should show what was recorded at each one and make the missing link obvious.

For a commercial site, I usually want identifiable AI referral sessions, engaged sessions, qualified leads or add-to-carts, completed conversions, assisted conversions where the analytics model supports them, and revenue. I annotate content releases and indexing events, but I do not call a coincident revenue change causal proof.

  • Define the qualified action before collection starts.
  • Use the authoritative order, lead, or booking record for completion.
  • Keep attribution-model output separate from directly observed events.

Build a report someone can act on

The report does not need to be complicated. It needs to make the next action obvious. I use separate columns for official citations, unique cited URLs, observed prompt citation rate, referral sessions, engagement, conversions, and revenue. Beside them I show sample sizes, dates, changes, and data limits.

If a custom score is useful for a long-running internal trend, I put it after the underlying measures and link the formula. The score is never the only output. A client should be able to disagree with the weighting and still use the raw evidence.

  • Show current value, previous comparable value, and the absolute change.
  • Annotate deployments, recrawls, indexing changes, campaigns, and major market events.
  • End with the decision, owner, due date, and the evidence that would reverse it.

The measurement setup I use

This model keeps each evidence source separate and joins records only when a shared identifier or documented rule supports the connection.

  1. Define the decision

    Choose the page, market, product, decision owner, comparison period, and action the result is meant to support.

  2. Store official citations

    Save platform, surface, cited URL, grounding query when available, citation count, filters, date range, and reporting limits.

  3. Run the frozen prompt panel

    Use three repeats per prompt, platform, locale, and device condition, then report observed rates with their denominators.

  4. Measure identifiable referrals

    Report sessions by raw source and medium, landing page, engagement, key events, and revenue without inventing missing referrals.

  5. Reconcile business records

    Confirm completed leads, orders, or bookings in the authoritative system and keep attribution assumptions visible.

  6. Publish the separate columns

    Show citation, observation, referral, conversion, and revenue measures side by side with samples, dates, changes, and limits.

What the report still cannot know

No available report covers every AI answer, user, platform, prompt, locale, and zero-click journey. Platform interfaces, definitions, sampling, and supported products can change after collection.

The separate ledgers improve honesty, but they do not create a causal bridge between exposure and revenue. That requires stronger experiments, clean identifiers, or both.

  • Google Search Console does not expose a documented AI-only performance segment.
  • Bing grounding queries are sampled and supported surfaces are not the whole AI market.
  • A prompt panel measures the selected panel rather than all real user demand.
  • Referral data can be stripped, misclassified, blocked, or absent in zero-click journeys.
  • Search Console and analytics APIs can apply limits, aggregation, or incomplete recent data.

Primary sources I checked

  1. AI features and your websiteGoogle Search Central, reviewed Opens in a new tab
  2. Introducing AI Performance in Bing Webmaster Tools Public PreviewMicrosoft Bing, reviewed Opens in a new tab
  3. The AI Performance dashboard and grounding query mappingMicrosoft Advertising, reviewed Opens in a new tab
  4. Publishers and Developers FAQOpenAI Help Center, reviewed Opens in a new tab
  5. Traffic acquisition dimensions and scopesGoogle Analytics Help, reviewed Opens in a new tab
  6. Search Analytics query APIGoogle Search Console API, reviewed Opens in a new tab

What changed

  1. Published the April research article with current Google, Bing, Microsoft, OpenAI, Search Console, and Analytics documentation.