Measurement playbook

How to Build an AI Visibility Tracking Baseline

Published By Piush Vaish

An AI visibility tracking baseline is a versioned record of how defined AI surfaces represent a company across a fixed set of buyer-relevant questions. To make that baseline useful for later comparison, freeze the prompts, product surfaces, material testing conditions, run design, observation fields, metric definitions and QA rules before T0. Retain the underlying answers and visible sources, then use the same measurement contract for later retests wherever possible.

The important word is not visibility. It is baseline.

A collection of screenshots tells you what happened. A baseline gives you a structured starting point against which later observations can be compared.

For B2B teams, that distinction matters because AI systems increasingly sit inside research, comparison and shortlist formation. A company may appear in an answer while being described inaccurately, placed in the wrong category, recommended for the wrong use case or framed differently from one AI surface to another.

Kojable uses AI answer alignment to describe the wider problem: reducing the gap between company reality, available public evidence and the story AI systems tell buyers. Visibility is one measurement dimension inside that wider problem.

This Guide shows how to establish the measurement layer before moving into diagnosis or improvement.

Start here

Build a versioned AI visibility tracking baseline that can support comparable retesting.

Goal
Build a versioned AI visibility tracking baseline that can support comparable retesting.
Inputs
An approved buyer-question panel, selected AI surfaces, a verified company truth set where accuracy is being measured, and defined collection conditions.
Output
A frozen T0 baseline, prompt panel, observation log, metric dictionary, source register, QA/version log and retest contract.

Baseline

A baseline is a measurement contract, not an audit

A baseline, tracking programme and audit perform different jobs.

Scroll horizontally if needed

Baseline, tracking, audit and benchmark comparison.
ActivityPrimary questionTypical outputWhen it happensWhat it does not prove
BaselineWhat is observable now under defined conditions?T0 measurement recordBefore improvement workWhy the answer occurred
TrackingWhat is changing across comparable observations?Time series and repeated observationsOngoingThat movement was caused by one intervention
AuditWhich gaps matter and what deserves investigation?Prioritised diagnosisAfter enough evidence existsExact hidden-model causality
BenchmarkHow does performance compare with another reference set?Relative comparisonWhen a suitable comparator existsThat the benchmark represents every buyer journey

A baseline therefore comes before diagnosis.

If a company is absent from a high-intent comparison answer, the baseline records that absence. If the company appears but is described as a mid-market tool when its current focus is enterprise, the baseline records that representation gap. If a competitor is recommended first, the baseline records the recommendation and surrounding evidence.

The baseline does not immediately declare what caused any of those outcomes.

That separation is central to Kojable's Monitor → Diagnose → Improve → Verify operating model. Monitoring establishes what was observed. Diagnosis interprets the meaningful gaps. Improvement defines the justified action. Verification later checks what changed.

For the broader monitoring framework, see Kojable's AI Brand Monitoring reference entry.

Questions

Start with buyer questions, not an arbitrary prompt count

The first input is not a number such as 20, 50 or 100 prompts. It is a set of buyer questions connected to actual decisions.

A B2B question panel may contain questions about category, fit, use case, comparison, integrations, pricing, proof, implementation, trust and alternatives.

The number of prompts should follow the decisions you need to observe, not a universal rule.

Kojable's 180-prompt Gemini study found that semantically related prompts tended to produce strongly related overall answers in that dataset. But the study does not show that nearby prompts always preserve brand mentions, citation sets, vendor order or recommendations. See How Many AI Prompts Do You Really Need to Track?.

Cluster related questions for efficiency, but preserve separate prompts when the difference could change a commercially meaningful outcome.

Scroll horizontally if needed

Prompt-variation treatment examples.
Prompt variationLikely treatment
“Best cash-flow tools for SaaS companies” vs “Best cash-flow software for SaaS”Candidate for the same family
“Best cash-flow tool for SaaS” vs “Best cash-flow tool that integrates with NetSuite”Keep separate if integration fit matters
“Best cybersecurity platform” vs “Best cybersecurity platform for banks in Ireland”Keep separate if geography or regulatory fit changes the decision
“[Brand] alternatives” vs “[Brand] vs [Competitor]”Different decision intent, keep separate

If your buyer questions have not yet been mapped, use the AEO Buyer-Question Mapping Playbook first. This Guide assumes that the important buyer decisions are already known.

Build a prompt panel

Every prompt should receive a stable identifier.

Scroll horizontally if needed

Example Prompt Panel with stable identifiers.
Prompt IDQuestion familyExact promptIntentPriorityPanel version
CAT-01CategoryWhat are the best [category] platforms for [use case]?DiscoveryHighv1
FIT-01FitWhich [category] tools are best for enterprise teams?FitHighv1
CMP-01ComparisonHow does [brand] compare with [competitor]?ComparisonHighv1
PRF-01ProofWhat evidence supports [brand]'s enterprise capabilities?ProofMediumv1

Do not quietly rewrite these prompts halfway through measurement. A meaningful wording change belongs in a new panel version.

Surfaces

Define the AI surfaces before collecting T0

“Test it in AI” is not a sufficient measurement definition.

Neither is “test ChatGPT, Gemini, Claude and Perplexity” if the product surface or mode changes the conditions being observed.

Kojable's cross-model citation study presented the same B2B buyer questions to Claude, Gemini, OpenAI and Perplexity and found that the visible evidence environments differed substantially under the study protocol. The study also explicitly warns that its one observed run per provider-question cell does not establish stable provider preferences or run-to-run behaviour. See Different Answers, Different Evidence.

The practical implication is not that every company must monitor four providers. It is:

Do not assume that observing one AI surface represents the entire AI-mediated buyer environment.

Choose the surfaces relevant to your buyers and record them precisely.

Useful fields include provider, product or surface, mode, model where visible or selectable, search or grounding state where controllable, and public-web versus connected-source scope.

Google states that AI features can use query fan-out, issuing multiple related searches across subtopics and data sources while developing a response. See Google Search Central guidance for AI features.

A baseline should therefore record the surface actually tested, not only the company that provides it.

Environment

Freeze the measurement environment

Repeating the same sentence is not automatically the same experiment.

Depending on the surface, relevant conditions may include language, geography, fresh versus existing conversation, account state, personalisation or memory, search state, source scope, selected model or mode, and date/time.

OpenAI documents that ChatGPT Search may use relevant saved memory when rewriting a search query and may use approximate location when returning local results. See Searching the web with ChatGPT.

Control what you reasonably can. Record what is observable and material. Never invent conditions the platform does not expose.

What to freeze before T0

Scroll horizontally if needed

Measurement-component freeze matrix.
Measurement componentFreeze before T0?WhyWhat to do if it changes
Exact promptYesWording can change the information needCreate a new panel/version
Prompt-panel versionYesKeeps question sets traceableRecord new version
Provider and product surfaceYesDifferent surfaces may behave differentlyTreat as a changed condition
Mode/search stateWhere controllableChanges available information pathsRecord separately
LanguageYesCan change content and available sourcesNew subgroup or version
GeographyWhere materialCan affect locally relevant answersRecord changed condition
Conversation stateYes where controllablePrior context may influence the answerStart fresh or flag contextual run
Run designYesAffects what the baseline representsNew measurement version
Metric definitionsYesChanging the denominator changes the metricVersion the formula
URL-normalisation ruleYesAffects source counts and overlapReprocess or version method
Company truth setYes for accuracy measuresAccuracy requires a defined referenceVersion when company reality changes
Raw answer retentionYesAllows later inspectionMissing evidence becomes a limitation

A useful baseline preserves both prompt reproducibility and environment reproducibility as far as the platform allows.

Runs

Match the run design to the decision

There is no evidence-backed universal answer to how many times every prompt should be run.

A single run can document a single dated observation. It cannot establish run-to-run stability.

Kojable's cross-provider study used one run per provider-question cell and explicitly states that repeated runs are required before the observed provider differences can be treated as stable traits. See Different Answers, Different Evidence.

Independent research reaches a compatible conclusion. A 2026 preprint by Ronald Sielinski found substantial variability in generative-search citation measurements and argues that single-run visibility results can create a misleading impression of precision. See Quantifying Uncertainty in AI Visibility.

Scroll horizontally if needed

Run-design implications by measurement objective.
Measurement objectiveRun implication
Capture a dated snapshotOne observation can be valid if labelled as such
Check whether an unexpected answer recursRepeated runs are useful
Estimate run-to-run variabilityMultiple comparable runs are required
Compare before and after an interventionPreserve the run design across checkpoints
Support a higher-stakes executive decisionRequire stronger evidence and disclose uncertainty

The IAB's 2026 AI visibility measurement framework similarly distinguishes directional measurement from decision-grade measurement and emphasises disclosure, stability and reproducibility. See Measuring Visibility in the AI Era.

Do not make a directional snapshot carry more certainty than its design supports.

Workbook

Build the Baseline Measurement Workbook

The workbook is the operational core of the baseline. It should expose the method rather than hide it behind a single visibility score.

Use the Baseline Measurement Workbook.

Use six logical layers.

1. Measurement Contract

Record objective, buyer decision being measured, T0 date, measurement version, prompt-panel version, provider/surface set, geography, language, run design, collection window and raw-evidence location.

2. Prompt Panel

Record prompt ID, exact prompt, family, buyer intent, branded or unbranded status, priority and panel version.

3. Observation Log

One row should normally represent one prompt × surface × run observation.

Illustrative row only. Replace placeholders with the actual collection record.

Scroll horizontally if needed

Illustrative Observation Log row.
FieldIllustrative value
Observation IDCAT-01-GEM-AIMODE-R01
Prompt IDCAT-01
ProviderGoogle
SurfaceAI Mode
Run1
TimestampYYYY-MM-DD HH:MM
LanguageEnglish
GeographyUnited Kingdom
Fresh conversationYes
Valid responseYes
Target company mentionedYes
Category accurateYes
Required capability presentNo
Competitor mentionedCompetitor A
Target recommendedNo
Visible citation presentYes
Target-domain citationNo
Factual issueNone observed
Raw response retainedYes

4. Metric Dictionary

Record metric name, question answered, numerator, denominator, analytical unit, eligible population, exclusions, deduplication, aggregation and material limitation.

5. Source Register

For each visible citation, retain observation ID, raw URL, normalised URL, domain, source type where useful, company-owned or third-party status, optional claim-support review and notes.

6. QA and Version Log

Record missing cells, failed calls, retries, methodology changes, reason for change, version introduced, comparability with prior version and T0 freeze status.

The objective is not to produce a prettier dashboard. It is to create an inspectable measurement record.

Metrics

Define the denominator before the percentage

A metric name without a counting rule is not a measurement specification.

Example metric dictionary

Scroll horizontally if needed

Example metric dictionary with definitions and important rules.
MetricDefinitionImportant rule
Brand mention rateEligible valid responses mentioning the company ÷ eligible valid responsesMultiple mentions in one response count once
Direct brand citation rateEligible citation-observable responses citing a verified company domain/page ÷ eligible citation-observable responsesFreeze the target-domain set
Recommendation rateEligible responses explicitly recommending the company ÷ eligible recommendation-intent responsesDefine recommendation before coding
Competitor co-mention rateEligible responses containing at least one frozen competitor ÷ eligible valid responsesKeep per-competitor flags separately
Required-attribute coveragePrompt-relevant required attributes represented accurately ÷ eligible required attributesDefine relevant attributes per prompt
URL source overlap|A ∩ B| ÷ |A ∪ B| for the canonicalised visible cited-URL sets of two comparable observationsDefine URL-level canonicalisation before T0
Citation volatility1 − mean pairwise Jaccard similarity of canonicalised visible cited-URL sets across comparable repeated runs for the same prompt × surface cellUse only with at least two eligible repeated runs

Kojable's current AI Brand Monitoring framework separates presence, representation, recommendations, competitors and evidence instead of collapsing them into one vague visibility score. See AI Brand Monitoring.

Do not calculate a metric such as “share of answer” until the numerator, denominator and counting unit are actually defined.

Citations

Record citations as evidence, not causal telemetry

Citations are useful baseline data because they are observable.

They can tell you whether visible sources were attached to an answer, which URLs appeared, whether company-owned sources appeared, which domains recur and whether source sets change between observations.

They do not automatically tell you that the source caused the answer, that the source was retrieved but not cited, that the platform considers it authoritative, that it was used in model training or that repeating the source will reproduce the answer.

Kojable's multi-platform citation-exposure research found that visible citation exposure varied materially by platform and query context. The study explicitly describes citation presence as an observable downstream signal rather than direct telemetry of a hidden retrieval process. See Retrieval Need and AI Citation Exposure Vary by Platform.

The cross-provider study separately measures citation events, claim placement and reviewed source support rather than treating them as the same thing. See Different Answers, Different Evidence.

That gives the workbook two different questions:

Was a citation observed?

and, where the team performs deeper review:

Does the cited page support the relevant claim?

Do not collapse them.

Method

Treat failures, retries and URL handling as part of the method

Failed observations, retries and URL normalisation affect what the baseline represents.

Failed observations

Illustrative arithmetic example: suppose the frozen panel contains 20 prompts, four surfaces and three planned runs. That creates 240 planned observation cells.

If eight calls fail, simply deleting those eight rows changes the denominator without documenting why.

Instead, retain the planned cell and classify it as valid response, failed response, blocked/refused response, timeout, capture failure, retry pending or excluded under a predeclared rule.

Kojable's cross-provider Research retained a failed provider-question cell and restricted the primary matched comparison to the complete panel rather than pretending every expected observation succeeded. See Different Answers, Different Evidence.

Retries

Define the retry policy before collection.

For example:

  • retry once after a technical failure;

  • do not replace a valid but commercially inconvenient answer;

  • retain both original and retry identifiers;

  • record which observation is eligible for the primary metric.

There is no need for one universal retry rule. There is a need for a documented one.

URL normalisation

The same source can appear as a clean canonical URL, a URL with tracking parameters, a fragment link, an alternate path or another URL that resolves to the same content.

At minimum retain both raw URL and normalised URL. Do not destroy the original evidence during cleaning.

Version

Version methodology changes instead of overwriting T0

Create a new measurement version when a change materially affects interpretation.

Typical version triggers include significant prompt rewrite, new or removed prompt family, different provider or product surface, changed language or geography, changed run design, changed collection mechanism, changed metric formula, changed eligibility rule, changed URL-normalisation method or a changed company truth set for accuracy scoring.

Illustrative version log:

Scroll horizontally if needed

Illustrative version log.
VersionDateChangeReasonComparable with prior version?
v1.0T0 dateInitial T0Baseline establishedN/A
v1.1Later checkpointAdded two validation promptsNew enterprise integration questionCore panel yes; new prompts have no historical T0
v2.0Later checkpointChanged product surfacePrevious surface retiredLimited comparability

Versioning does not make the data perfect. It makes the limitations visible.

QA

Check whether the baseline is fit for purpose

Before anyone turns the results into an action plan, perform a measurement QA.

A useful baseline should answer five questions:

The IAB's framework separates directional measurements from measurements intended to support stronger decisions and calls for greater disclosure around stability and reproducibility. See Measuring Visibility in the AI Era.

The goal is to know what the evidence can and cannot support.

Retest

Freeze the retest protocol before making improvements

Do not decide how you will measure improvement after seeing the new answer.

Define the retest contract at T0. Record which prompts will be rerun, which surfaces will be tested, run count, geography and language, metric definitions, eligibility rules, source-normalisation method, comparison checkpoints and how unavoidable platform changes will be recorded.

When the team later publishes new evidence, updates positioning, fixes a third-party profile or changes owned content, run the comparable panel again.

You may observe that the company is now mentioned where previously absent, category framing changed, a required capability is now included, a competitor is no longer recommended first, the citation set changed, a company-owned source appeared, source overlap increased or decreased, or the result remained unchanged.

Those are observations. They do not automatically prove that one intervention caused the change.

Example

Worked example: a B2B company changing its market position

Consider an illustrative B2B software company moving from a mid-market positioning towards enterprise buyers. This is not a customer result.

Step 1: define the decision

Measurement objective: establish whether selected AI surfaces describe the company as appropriate for enterprise buyers across relevant category, fit and comparison questions.

Step 2: freeze the question panel

Scroll horizontally if needed

Worked-example prompt panel.
IDPrompt familyPrompt
CAT-01CategoryWhat are the leading [category] platforms for enterprise companies?
FIT-01FitIs [company] suitable for large enterprise teams?
CMP-01ComparisonHow does [company] compare with [competitor] for enterprise use?
PRF-01ProofWhat evidence supports [company] for enterprise deployments?

Step 3: define the environments

Record selected provider surfaces, English, defined geography, fresh conversations, search state where controllable and planned run design.

Step 4: define required attributes

For FIT-01, the truth set might include current target audience, enterprise capability, security/compliance facts, deployment requirements and relevant integration facts.

Only verified current company facts belong in that truth set.

Step 5: collect T0

Suppose one illustrative observation describes the company using its older mid-market positioning.

The baseline records the exact answer, incorrect audience framing, competitors named, visible citations, source URLs, environment and run.

It does not yet say:

“Directory X caused the old description.”

The diagnosis stage can later inspect whether cited or recurring sources themselves contain older positioning and whether other evidence gaps exist.

Step 6: freeze the retest contract

Before changes are made, record exactly how the same question panel will be retested.

That is what turns the initial collection into a baseline rather than a temporary screenshot archive.

Sources and further reading

Frequently asked questions

How many prompts should an AI visibility baseline include?

There is no universal prompt count. Start with the buyer decisions that matter to the company, cluster genuinely similar questions where useful, and retain separate validation prompts when differences in persona, geography, integration, competitor or commercial intent could change the answer.

How many times should each AI prompt be run?

One run records one snapshot. Repeated comparable runs are needed when you want to assess run-to-run variability or support stronger comparative conclusions. Current research does not establish one universally correct number of runs for every platform, topic and decision.

Which AI platforms should a company include?

Include the provider and product surfaces relevant to the buyer journey you are measuring. Do not assume one provider represents every AI environment, but do not add platforms solely to reach an arbitrary minimum either.

Can I build an AI visibility tracking baseline in a spreadsheet?

Yes. A spreadsheet can work well if it preserves the measurement contract, exact prompts, observation-level records, raw evidence, metric definitions, citations, failure states and methodology versions.

What is the difference between an AI visibility baseline and an AI visibility audit?

A baseline establishes what was observed at T0 under defined conditions. An audit interprets those observations to identify which representation, competitor, source or evidence gaps deserve investigation. The baseline is the measurement input. The audit is a diagnostic process applied to that evidence.

From baseline to diagnosis

Workbook output

A well-built baseline should leave you with evidence, not conclusions disguised as evidence.

You should know which buyer questions were tested, which AI surfaces were observed, under what conditions, what the systems said, what they omitted, which competitors appeared, which visible sources were attached, how each metric was calculated, which observations failed and how the measurement will be repeated.

That is the end of the baseline stage.

The next question is different:

Which observed gaps are commercially meaningful, what evidence is associated with them and what deserves investigation first?

Once your T0 baseline is frozen, continue to the Competitive Gap Audit Guide to move from Monitor into Diagnose.

About this guide

Kojable is an AI answer alignment platform for B2B companies. It helps teams monitor how AI represents a company, diagnose the source and information gaps associated with those answers, guide practical improvements, and retest comparable questions to verify what changed.

KojableAI Answer Alignment

View all guides