Article / A straightforward way to check the claim
How I Verify What Actually Works in AI Search
A citation, crawler hit, visit, and sale prove different things. This is how I keep AI search sources, tests, interpretation, and results honest.
When somebody tells me a page is winning in AI search, my first question is simple: what exactly happened? A crawler visit, a citation, a click, and a sale are all useful, but they are not interchangeable. This is how I keep the claim tied to the thing that was actually observed.
The rules I keep coming back to
- Start with the decision, then choose the measure that can actually answer it.
- Keep eligibility, exposure, citation, visit, conversion, and revenue separate.
- Save the test conditions, repeat the test, and publish the limits beside the finding.
Quick decision table
Use the proof that matches the claim
One screenshot cannot prove the whole journey. Each outcome has a different place where it can be checked and a clear point where the conclusion has to stop.
| What happened | Where I check it | What it does not prove |
|---|---|---|
| A crawler reached the page | Server logs, HTTP response, robots rules, CDN decision, rendered HTML, and canonical state. | That the page was indexed, retrieved for an answer, cited, or preferred. |
| The page appeared in an AI answer | The full answer, exact prompt, visible source URL, product surface, locale, account state, and timestamp. | A stable rank, broad market visibility, endorsement, or future recurrence. |
| A platform reported citations | The platform report with its date range, dimensions, filters, supported surfaces, and stated reporting limits. | Answer placement, authority, referral traffic, or commercial value. |
| A visit came from an AI product | Session-level analytics, source and medium, landing page, engagement, and the next recorded action. | That a specific citation caused the visit or that the visitor converted. |
| An agent completed a task | The final interface state plus the authoritative receipt in the order, booking, form, or publishing system. | That every agent, device, account, or version can repeat the task safely. |
Start with the decision, not the dashboard
Before I open a report, I write down the decision. Are we trying to repair crawler access, improve a source page, fix product data, understand a drop in referrals, or test whether an agent can finish a task? If the decision is vague, the report will become a pile of numbers looking for a story.
Then I choose one primary measure that can change that decision. Citation recurrence can help decide whether a source page deserves work. Qualified sessions and revenue answer a commercial question. Neither one replaces the other, and a blended score usually hides that distinction.
- Name the person who will act on the result.
- Write the decision date and the pages, market, and products in scope.
- Choose one primary measure and keep diagnostic measures clearly secondary.
Keep the measurement layers separate
I keep technical eligibility, answer exposure, citation or brand mention, referred visits, assisted conversions, and completed business outcomes separate. A page can improve at one layer and do absolutely nothing at the next. That is normal. It only becomes a problem when the report pretends the layers are one funnel with perfect tracking.
Official documentation tells us how a platform says a feature works. Property data tells us what the platform recorded for that site. A prompt test is one bounded observation. A business record tells us what completed. Any connection between those records is an interpretation unless the system exposes a direct link.
- Label platform statements, site data, observations, and interpretations separately.
- Keep the original dimensions and filters attached to exported data.
- Do not present correlation as a causal chain.
Save the environment behind every answer
Generated answers can change by product surface, time, location, language, device, account state, memory, personalization, model, and prompt wording. A screenshot without those conditions is decoration. It is not a reproducible record.
I save the exact prompt, timestamp and time zone, country, language, device class, signed-in state, and the product or model label visible in the interface. I also save the complete answer, the source URLs, and the specific claim each source appears to support.
- Keep the verbatim prompt and any declared paraphrases.
- Record destination URLs, not only source domains.
- Capture errors, omissions, and wrong brand facts with favorable mentions.
Repeat the test before you call it a pattern
One generated response is one observation. It is not a rank. I set the prompt panel, platforms, repeat count, collection dates, and controlled conditions before I look at the result. That makes it much harder to quietly discard the runs that ruin a good story.
For each prompt family, I report citation incidence, recurring URLs, factual accuracy, and the range across runs. I keep each engine and product surface separate because they expose different data and can use different retrieval and answer systems.
- Declare the sample before reviewing the result.
- Include failed runs and missing answers.
- Show the numerator and denominator behind every reported rate.
Connect visibility to the business without forcing the connection
AI visibility can influence discovery without producing an immediate click. It is worth tracking, but it is not the finish line. I keep referral sessions, engaged sessions, qualified actions, assisted conversions, completed sales, and revenue in separate columns.
When a referral tag exists, I connect the session to the landing page and next action. When it does not, I say so. Branded search, direct visits, and sales changes can add context. They do not automatically belong to an AI citation just because the dates look convenient.
For agent actions, verify the final receipt
Crawler access is only the first gate. A browser agent also has to understand the controls, preserve state, handle errors, request approval, and reach a clear final state. Reaching the checkout or filling a form is not completion.
For any action that matters, I confirm the result in the authoritative business system. That can be the order record, booking, sent form, published page, or payment receipt. It is the same rule I use everywhere else: the proof surface must belong to the outcome.
- Repeat the same meaningful task more than once.
- Capture the exact failure point and interface state.
- Check the final action outside the agent whenever a business record exists.
Method
How I run the review
The size of the project can change. The evidence chain should not. This is the shortest version of the process that still lets another person audit the conclusion.
Frame the decision
Name the owner, scope, decision date, primary measure, and the action each possible result would trigger.
Check the official source
Use current first-party documentation for platform behavior and record the date every source was reviewed.
Check the real delivery
Inspect the response, crawl controls, rendered content, canonical state, and machine-readable data on the live route.
Collect controlled observations
Preserve prompts, conditions, answers, source URLs, errors, and repeated samples without deleting inconvenient results.
Reconcile the layers
Keep platform exposure, citation, referral, conversion, and revenue separate before interpreting how they may relate.
Publish the limits
State what the records cannot establish, show the real dates, and update the history after a material change.
Limits
What this method cannot prove
This method makes a claim easier to audit. It cannot reveal private ranking or retrieval systems, and it cannot guarantee crawling, indexing, citation, recommendation, traffic, or revenue.
Platform reports measure different things and can change their availability, aggregation, and definitions. Every finding applies to the recorded conditions and dates, not to every possible user experience.
- Crawler access establishes possibility, not inclusion or preference.
- Repeated prompts describe the sampled surfaces, not the entire market.
- A citation does not establish endorsement, placement, influence, or commercial value.
- A before-and-after result does not prove causation when other conditions also changed.
- One successful agent task does not prove universal reliability or safety.
Sources
Primary sources I checked
- AI features and your websiteGoogle Search Central, reviewed Opens in a new tab
- Introducing AI Performance in Bing Webmaster Tools Public PreviewMicrosoft Bing, reviewed Opens in a new tab
- Publishers and Developers FAQOpenAI Help Center, reviewed Opens in a new tab
- Build agent-friendly websitesweb.dev, reviewed Opens in a new tab
- Article structured dataGoogle Search Central, reviewed Opens in a new tab
Updates
What changed
Published this article, checked every linked primary source, and separated publication history from the February research month.
