Measurement case study · AI search visibility

Same Visibility Metric, Different Problem: What Two B2B Baselines Revealed

AI search visibility is useful as a baseline, not a diagnosis. In two anonymised US-English B2B benchmarks, Company A appeared in 4 of 164 usable unprompted answers, while Company B appeared in 56 of 105. The panels were different and are not a head-to-head comparison. Their value was diagnostic: Company A showed low answer-level presence across its tracked buyer questions, while some Company B intents were already at 15 of 15 mentions. Those different states led Kojable to develop a staged measurement framework spanning source uptake, information reflection, answer quality, robustness and commercial validation.

Case question: How should a B2B company interpret an AI search visibility baseline, and what should it measure next?

Published By Kojable
Programmes
2 anonymised B2B benchmarks
Planned checks
330
Surfaces
5 AI/search surfaces
Repetitions
3 per question/surface

Measurement canvas

From buyer questions to the next measurement decision

The important finding was not that one company had a bigger percentage than the other. It was that the same type of visibility metric exposed two different measurement problems.

Benchmarks
2
Planned checks
330
Surfaces
5
Repetitions
3
  1. 01

    Buyer questions

    The benchmark used deliberately selected buyer questions rather than attempting to estimate every prompt a buyer might ask.

  2. 02

    Repeated observations

    Each question was measured three times across five user-facing AI/search surfaces.

  3. 03

    Unprompted discovery + citations

    Record whether the company appeared and which sources were visibly cited.

  4. 04

    Diagnosis + next measurement

    The next metric should depend on the diagnosed problem.

Interpretation boundary

The baseline describes the observed answer-layer starting state. It does not by itself diagnose the hidden upstream cause or establish intervention causality.

Programme context

Kojable developed the measurement methodology and conducted the AI-answer and citation analysis as the research and intelligence layer of an anonymised client programme. The methodology covers prompt design, source discovery, extraction, scoring, orchestration, validation and quality control.

The underlying measurement records retain the exact questions, answer text, citations, AI/search surface, repetition number and machine-recorded timestamps.

The important finding was not that one company had a bigger percentage than the other. It was that the same type of visibility metric exposed two different measurement problems.

That changed the question from:

How visible is the company?

to:

What does this starting state tell us to investigate next, how should that change be measured, and what evidence would justify a stronger conclusion?

Baseline design

How was AI search visibility measured?

The benchmark used deliberately selected buyer questions rather than attempting to estimate every prompt a buyer might ask.

Each question was measured three times across five user-facing AI/search surfaces:

  • ChatGPT
  • Gemini
  • Google AI Overview
  • Google AI Mode
  • Claude

The study was configured for the United States market and English-language responses.

Measurement fields for the two anonymised B2B benchmarks
Measurement fieldCompany ACompany B
Buyer questions1210
AI/search surfaces55
Repetitions per question/surface33
Planned observations180150
Collection date10 Sep 202611 Sep 2026
MarketUnited StatesUnited States
LanguageEnglishEnglish
Headline discovery metricUnprompted Brand Discovery RateUnprompted Brand Discovery Rate

Each repetition remained a separate observation rather than being averaged into a single generated answer.

That matters because generative outputs vary. OpenAI's current evaluation guidance explicitly notes that models can produce different outputs from the same input, recommends task-specific and continuous evaluation, and advises combining numerical metrics with human judgement.

Three repetitions were the design used in this benchmark. They should not be treated as a universal optimum for every AI visibility study.

Measurement contract

For this case study, the visibility percentages should be interpreted only with their measurement contract:

questions → eligibility → numerator → denominator → surfaces → repetitions → market/language → collection window → failure handling

Without those elements, a percentage labelled “AI visibility” can represent materially different things.

A technical collection failure was excluded from answer-level outcome metrics. A valid observation in which an AI Overview was not shown was retained as a valid surface outcome and treated differently from a collection failure.

What is Unprompted Brand Discovery Rate?

Unprompted Brand Discovery Rate is the share of usable answers to eligible buyer questions that do not name the company in which the company nevertheless appears.

Formula

eligible unbranded answers naming the company ÷ usable eligible unbranded answers

The analytical unit is one recorded answer observation.

Questions explicitly naming the client remain useful for analysing known-brand representation, but they are excluded from the unprompted buyer-discovery denominator.

That distinction prevents visibility from being inflated simply because the user supplied the company name.

A question asking:

“What does Company X offer?”

answers a different question from:

“Which providers can solve this problem?”

The first tests representation once the company is already known. The second tests whether the company enters the buyer's answer without being supplied by the user.

Related Kojable research, Branded AI Visibility Is Not Market Visibility, examines the same denominator problem from a different dataset.

Company A baseline

What did Company A's baseline show?

Company A's panel contained 11 unprompted buyer questions after one explicitly branded question was separated from the discovery metric.

That produced:

11 questions × 5 surfaces × 3 repetitions = 165 planned unprompted observations

One Google AI Overview collection failed technically.

It was excluded rather than scored as a negative brand outcome, leaving 164 usable unprompted observations.

Company A appeared in 4 of those 164 answers, or 2.4%.

Company A unprompted brand discovery by AI/search surface
SurfaceBrand mentionsUsable unprompted answersDiscovery rate
ChatGPT0330.0%
Gemini0330.0%
Google AI Overview1323.1%
Google AI Mode1333.0%
Claude2336.1%
Total41642.4%

Figure 1 — Company A: unprompted brand discovery by AI/search surface

Figure annotation: 4 of 164 usable answers · 2.4%

Horizontal bars on a fixed 0–100% scale show rates of 0.0%, 0.0%, 3.1%, 3.0% and 6.1%. 0%25%50%75%100% ChatGPT0.0% Gemini0.0% Google AI Overview3.1% Google AI Mode3.0% Claude6.1%
Company A had low unprompted answer-level presence across the September 2026 US-English benchmark. One technical Google AI Overview failure was excluded rather than scored as a negative result. The chart describes the observed answer-layer state, not its hidden upstream cause. Scale: 0–100%.

Source: Kojable anonymised AI-answer measurement programme, 10 September 2026.

The defensible finding is:

Company A had low unprompted answer-level presence across its selected buyer-question benchmark.

It does not tell us why.

A company can be absent from an answer because of discoverability, retrieval, selection, synthesis, the available evidence environment or another factor that cannot be established from the endpoint alone.

The result identifies where the problem is visible at the answer layer; it does not establish whether the upstream constraint is discoverability, retrieval, selection, synthesis or the public evidence environment.

So the baseline identifies where the problem is visible, not automatically what caused it.

Company B baseline

What did Company B's baseline show?

Company B contained seven unprompted buyer questions after three explicitly brand-named questions were separated.

That produced:

7 questions × 5 surfaces × 3 repetitions = 105 unprompted observations

Company B appeared in 56 of 105 answers, or 53.3%.

Company B unprompted brand discovery by AI/search surface
SurfaceBrand mentionsUnprompted answersDiscovery rate
ChatGPT72133.3%
Gemini132161.9%
Google AI Overview142166.7%
Google AI Mode122157.1%
Claude102147.6%
Total5610553.3%

Figure 2 — Company B: unprompted brand discovery by AI/search surface

Figure annotation: 56 of 105 answers · 53.3%

Horizontal bars on a fixed 0–100% scale show rates of 33.3%, 61.9%, 66.7%, 57.1% and 47.6%. 0%25%50%75%100% ChatGPT33.3% Gemini61.9% Google AI Overview66.7% Google AI Mode57.1% Claude47.6%
Company B showed stronger answer-level presence inside its own September 2026 US-English benchmark, including some intent cells at 15 of 15 mentions. Company A and Company B used different question panels, so the percentages are not a head-to-head ranking. Scale: 0–100%.

Source: Kojable anonymised AI-answer measurement programme, 11 September 2026.

The more useful finding came from looking below the overall percentage.

Two Company B buyer-intent areas produced mentions in 15 of 15 measured answers. Another tracked intent produced none.

That changes the measurement problem.

If a company already appears in 15 of 15 answers for an intent, raw mention visibility has no remaining headroom inside that benchmark.

The next question should not be:

How do we increase the mention rate?

It should become:

Is the company represented accurately, differentiated appropriately, supported by useful evidence and recommended where it genuinely fits?

Interpretation

Why didn't 2.4% and 53.3% mean the same thing?

Because a visibility percentage describes an observed state.

It does not diagnose that state.

How an observed AI visibility state changes the next measurement question
Observed stateWhat the result tells usWhat it does not establishMore useful next question
Low unprompted presenceThe company rarely enters the measured buyer answersWhy the company is absentWhich evidence, entity or information gap needs investigation?
High or ceiling-level presenceBrand inclusion is already strong for that intentThat representation quality is goodIs the company described accurately and distinctively?
Citation without visible brandA source was attributedThat the brand became visibleWhat did the source actually support?
Brand mention without target citationThe company entered the answerThat the intended page was usedWhich evidence paths are associated with the answer?
Strong presence with unknown recommendation qualityThe brand is visibleThat it is an appropriate recommendationDoes the company genuinely fit the buyer's requirements?
One post-change increaseThe observed result movedThat the intervention caused itDoes movement persist against a credible comparison?

This is the central finding of the case study:

Answer

A visibility metric establishes the starting state. The next metric should depend on the diagnosed problem.

Kojable is an AI answer alignment platform for B2B companies. Visibility is one measurement dimension inside the wider operating model:

Monitor → Diagnose → Improve → Verify.

Measurement gap

Why were visibility and citation counts not enough?

The original programme could answer several useful questions:

  • Did the company appear?
  • On which AI/search surface?
  • For which buyer question?
  • Which sources were visibly cited?

But analysing those results exposed questions those metrics could not answer on their own.

A brand-related citation did not necessarily mean the intended intervention page appeared.

A cited page did not necessarily mean the important information on that page entered the generated answer.

A mention did not necessarily constitute a recommendation.

A recommendation would not necessarily be useful if the company did not objectively fit the buyer's requirements.

And an increase in AI visibility would not establish commercial incrementality.

The measurement system therefore needed to evolve from:

Did the company appear and how many citations were returned?

towards:

Did the intended source surface, did the changed information enter the answer, did representation quality remain acceptable, did the result persist and did any downstream buyer behaviour change?

Expanded framework

The measurement system that emerged from the baseline work

Rather than turning one visibility score into a larger composite score, Kojable separated the programme into distinct measurement questions.

The staged AI-answer measurement system developed from the baseline analysis
LayerMetricQuestion answered
Buyer discoveryUnprompted Brand Discovery RateDoes the company enter relevant buyer answers without being named?
Source uptakeTarget-Source Uptake RateDoes the exact intervention URL begin appearing?
Information reflectionChanged-Information Reflection RateDoes the introduced or corrected information actually enter the answer?
Evidence qualityCitation Support RateDoes the cited source genuinely support the associated claim?
Representation qualityAnswer AccuracyAre material factual claims correct?
Answer CompletenessAre the required decision-relevant facts included?
Recommendation qualityAppropriate Recommendation RateIs the company recommended where it objectively fits?
Inappropriate / Misrepresentation RateIs optimisation producing unsuitable or inaccurate recommendations?
Recommendation Share of VoiceHow frequently is the company recommended relative to named alternatives?
DurabilityPersistenceDoes the effect hold across later periods?
Prompt RobustnessDoes it survive semantically equivalent prompt variants?
Cross-System RobustnessDoes the result hold across more than one AI/search surface?
System-Transition RetentionDoes the effect survive a material model/product transition?
Latency to UptakeHow long does sustained source or information uptake take?
Observed market reachGenerative-AI observed reachIs the site appearing in real-world Google generative-search exposure?
Qualified acquisitionAI Assistant qualified valueAre AI-assistant visits producing useful buyer actions?
Commercial validationPipeline/revenue/contribution treatment effectDid an intervention create incremental business value?

Only the baseline discovery and original citation environment were historically measured in these two programmes.

The later measures are the prospective framework developed from the analysis, not retrospective results.

Figure 3 — From discovery to commercial validation
  1. Historically measured

    Measured baseline

    Unprompted buyer discovery

    Baseline citation environment

  2. Prospective

    Source and answer measurement

    Target-source uptake → Changed-information reflection → Citation support → Accuracy / completeness → Appropriate recommendation

  3. Prospective

    Durability

    Persistence → Prompt robustness → Cross-system robustness → System-transition retention

  4. Prospective

    Downstream validation

    Observed reach → Qualified buyer action → Commercial validation

Evidence boundary

Only Unprompted Brand Discovery and the original citation environment were measured historically in these baselines. Later stages are the prospective measurement framework developed from the analysis.

This is a measurement architecture. It does not imply that every commercial AI product exposes each hidden processing stage.

What is Target-Source Uptake Rate?

General citation volume is useful for understanding the source environment.

It becomes less specific when the research question concerns a particular intervention.

Target-Source Uptake Rate is the share of eligible answers that cite the exact URL affected by an intervention.

Formula

eligible answers citing the exact intervention URL ÷ eligible answers

Suppose a team publishes a new research page to clarify an outdated product fact.

If another company page starts appearing more frequently, total first-party citation activity may rise without showing whether the intended research page has begun surfacing.

Target-source measurement keeps the outcome connected to the specific asset being tested.

It should therefore be registered alongside:

  • intervention URL;
  • deployment date;
  • affected buyer questions;
  • changed claim or evidence;
  • crawl/index confirmation where applicable;
  • post-intervention observation window.

Target-Source Uptake Rate was not imposed retrospectively on these baselines. It is a prospective intervention metric.

What is Changed-Information Reflection Rate?

A target page being cited still does not answer the most important information question:

Did the changed information enter the answer?

Kojable therefore defines Changed-Information Reflection Rate as:

The share of eligible answers that correctly express a pre-specified fact introduced or corrected through an intervention.

Formula

eligible answers correctly expressing the changed fact ÷ eligible answers

The target fact should be defined before deployment.

For example, if a page contains an ambiguous capability statement and the intervention adds a verified clarification, measurement should not stop at whether the page appears as a citation.

It should also ask:

Does the subsequent answer correctly reflect the clarification?

This lets the programme distinguish two materially different results.

Target citation rises, reflection does not

The page may be surfacing, but the intended information is not consistently appearing in answers.

That suggests a need to investigate relevance, clarity, source context or other information pathways.

Reflection rises, target citation does not

The changed information may be reaching answers through another public source or without visible attribution to the intended page.

That warrants investigation rather than a claim of direct citation causality.

Neither outcome alone proves the hidden causal path, but the distinction provides substantially more diagnostic value than citation volume alone.

Why use a more specific citation taxonomy?

A generic “brand-related citation” flag becomes insufficient once the programme starts measuring interventions.

Different source relationships mean different things.

Kojable's expanded framework separates:

Citation classes used by the expanded measurement framework
Citation classMeaning
Exact target URLThe specific page affected by the intervention
First-party brand citationA citation on the company's owned domain
Explicit third-party brand citationA third-party title, source or URL explicitly identifying the company
Content-verified brand citationA third-party page that materially discusses the company even when the title or URL does not
Competitor/comparison citationA comparison or competitor source discussing the company
Supporting non-brand sourceA source supporting the category or answer without materially discussing the company

This avoids treating all citations as interchangeable.

It also avoids several common inference errors.

A third-party comparison citation is not the same event as an exact first-party target page surfacing.

A supporting category source does not necessarily create company visibility.

And a visible citation does not, by itself, prove that the source caused the answer.

Historical series should not be silently rewritten when citation definitions change. Where classifications are improved, retained evidence can be rescored explicitly or old and new definitions can be reported as separate series.

Are mentions, citations and recommendations the same metric?

No.

Mention ≠ citation ≠ recommendation.

A mention records that the company appeared in the visible answer.

A citation records that the surface visibly attributed a source.

A recommendation requires a separate judgement that the answer is presenting the company as an appropriate option for the buyer's stated need.

The original baseline measured mentions and citation behaviour.

It did not contain a separate historical recommendation rubric, so this case study does not retrospectively create a recommendation percentage.

That is important because an optimisation programme should not reward a system simply for recommending the company more often.

A stronger future measure needs an eligibility rule.

For example, the programme might define eligibility using verified:

  • use cases;
  • product capabilities;
  • company size;
  • deployment requirements;
  • integrations;
  • security requirements;
  • pricing constraints.

The result can then be classified as an appropriate recommendation only where the company genuinely fits.

How should answer quality be protected?

More visibility is not automatically better answer alignment.

A company can become more visible while being represented incorrectly.

A recommendation can increase while becoming less appropriate.

An intervention can improve a visibility proxy while reducing answer quality.

The framework therefore introduces quality guardrails.

Citation Support Rate

Checked cited claims genuinely supported by their cited source ÷ cited claims checked

This asks whether the visible source actually supports the associated statement.

It does not claim to reveal the system's hidden reasoning.

Answer Accuracy

Correct material factual claims ÷ material claims scored

The material claims and scoring standard must be defined before evaluation.

Answer Completeness

Required decision-relevant facts correctly covered ÷ pre-registered required facts

This catches omissions as well as outright factual errors.

Appropriate Recommendation Rate

Eligible buyer answers in which the company is appropriately recommended ÷ eligible recommendation-intent answers

The eligibility rule must be fixed before collection.

Inappropriate Recommendation / Misrepresentation Rate

This is the guardrail.

If recommendation visibility increases while factual accuracy or buyer fit decreases, that is harmful optimisation, not success.

Durability

How do you know whether an improvement is durable?

A movement appearing once on one surface for one wording is a weaker result than one that survives time, prompt variation and system changes.

The framework therefore separates the size of an initial movement from its durability.

Durability measures and the questions they answer
MeasureQuestion answered
PersistenceDoes the result continue across later measurement periods?
Prompt RobustnessDoes it survive semantically equivalent buyer questions?
Cross-System RobustnessDoes the effect appear in the intended direction across more than one measured surface?
System-Transition RetentionWhat happens when a model or product materially changes?
Latency to UptakeHow long does it take from deployment or index confirmation to sustained source/information uptake?

Instead of attempting to predict exactly what an undisclosed future model will prefer, the programme can ask:

Which interventions continue to produce useful, accurate effects across prompts, systems, time and model/product transitions?

That is measurable resilience rather than speculative future-model optimisation.

Evidence strength

When can an AI visibility change support a causal claim?

A post-intervention increase is an observation.

Timing alone does not establish cause.

Kojable therefore separates evidence states.

Figure 4 — Evidence strength after an intervention
  1. Observed

    A post-change association appears.

  2. Supported

    The exact source and/or changed information repeatedly appears.

  3. Controlled

    A credible treatment/comparison design shows differential movement.

  4. Causal

    A randomised or sufficiently strong quasi-experimental design supports a treatment-effect claim.

  5. Commercially causal

    The downstream business outcome also survives an appropriate counterfactual analysis.

Historical boundary

The two historical baseline programmes did not progress through or achieve every evidence level.

This helps match claim strength to evidence strength.

A target URL beginning to appear after deployment is more informative than total citation volume.

A target URL appearing alongside correct reflection of a newly introduced fact is stronger again.

A credible treatment/control design supports stronger intervention inference.

Commercial incrementality needs its own downstream comparison.

The framework does not require teams to stay silent until a perfect experiment exists. It requires them to describe the evidence they actually have.

What should different post-intervention patterns mean?

The framework turns combinations of measures into decisions.

Prospective post-intervention decision rules
Observed patternInterpretationNext action
Target citation ↑, reflection unchangedThe source is surfacing, but the intended information is not consistently entering answersExamine content clarity, relevance and citation context
Target citation unchanged, reflection ↑Information may be entering through another evidence pathInvestigate third-party propagation and other sources
Citation ↑, reflection ↑, appropriate recommendation ↑ and persistsStrong answer-layer evidenceReplicate on another topic cluster and test commercial linkage
AI-answer metrics ↑, commercial outcomes unchangedAnswer-layer movement is established; business value remains unprovenContinue the commercial lag window and inspect intent/conversion quality
AI-answer and commercial treatment effects ↑ against credible controlsStrongest evidence stateConsider scaling while retaining holdouts or controls
Single-period jumpInsufficient durability evidenceRepeat before drawing a strategic conclusion
Recommendation ↑, accuracy or fit ↓Harmful optimisationCorrect or roll back the intervention

These are future decision rules, not retrospective claims about Company A or Company B.

Reporting cadence

How should reporting change as the programme matures?

Not every metric matures on the same timescale.

A weekly AI-answer report should not pretend to prove quarterly revenue incrementality.

A quarterly business review should not be dominated by screenshots from individual prompts.

The framework therefore separates reporting by decision horizon.

Reporting cadence by decision horizon
CadencePrimary jobExample measures
WeeklyDetect answer-level movement and operational issuesDiscovery, target-source uptake, reflection, accuracy, citation support, recommendation fit, unavailable runs, model/system anomalies
MonthlyConnect answer movement with observed reach and qualified demandGoogle generative-AI impressions, branded/non-branded demand, Organic Search, AI Assistant traffic, qualified visits, trials/demos/leads
QuarterlyEvaluate commercial and causal evidenceTreatment/control effects, opportunities, pipeline, new ARR/bookings, revenue, contribution, cohort outcomes
Figure 5 — Reporting cadence

Weekly

Discovery, target-source uptake, reflection, accuracy, citation support, recommendation fit, unavailable runs, system/model anomalies.

Monthly

Generative-AI impressions, branded/non-branded demand, Organic Search, AI Assistant traffic, qualified visits, trials/demos/leads.

Quarterly

Treatment/control effects, opportunities, pipeline, new ARR/bookings, revenue, contribution, cohort outcomes.

Decision horizon

Answer-layer movement, observed reach and commercial incrementality remain separate reporting layers.

Google's Search Console now provides dedicated generative-AI performance reporting, including impressions plus page, country, device and date dimensions. Google says the reports had rolled out worldwide by 31 August 2026.

That observational Google layer should complement, not replace, the controlled buyer-question benchmark.

The two populations are different.

The benchmark exposes the actual answer, brand presence and citations for a fixed question panel.

Search Console shows how URLs from the site are appearing within Google's generative Search features in real usage.

GA4 now also defines an AI Assistant channel for sources including ChatGPT and Gemini. Google's AI Overviews and AI Mode are excluded from that channel and included within Organic Search. See Google's Default channel group documentation.

That enables a future measurement chain such as:

controlled answer benchmark → observed generative reach → AI/organic acquisition → qualified action → pipeline/revenue

The stages should still be reported separately.

Why does the measurement contract matter?

AI visibility is becoming an industry measurement category, but methodologies still differ substantially.

IAB's August 2026 Measuring Visibility in the AI Era guidance was developed in response to that problem. IAB says more than 20 companies now sell AI visibility measurement tools and that differing methodologies can produce different answers for the same brand or publisher.

That supports an important conclusion from the Kojable programmes:

An AI visibility percentage has limited meaning without its measurement contract.

At minimum, the reader needs to know:

  • Question panel

    which questions were measured;

  • Eligibility

    which answers were eligible;

  • Numerator

    the numerator;

  • Denominator

    the denominator;

  • Surfaces

    which AI/search surfaces were included;

  • Repetitions

    how many repetitions were run;

  • Market and language

    market and language;

  • Collection dates

    collection dates;

  • Failure handling

    how unavailable observations were handled.

That is also why Company A's 2.4% and Company B's 53.3% cannot be turned into a direct brand ranking.

Both percentages can be correctly calculated while answering different benchmark questions.

Operating sequence

How does the measurement programme evolve after the baseline?

The analysis produced not only new metrics but an operating sequence.

Measurement programme phases, purposes and outputs
PhasePurposeOutput
FoundationFreeze prompts, preserve raw responses and citations, standardise definitionsReproducible baseline
CalibrationEstimate ordinary variation and establish scoring rubricsVariance and scoring model
ExperimentRegister changed claims, intervention URLs and comparisonsTestable intervention hypothesis
Post-intervention measurementMeasure target-source uptake and changed-information reflectionAnswer-layer evidence
RobustnessTest persistence, paraphrases and cross-system behaviourDurability evidence
Commercial linkageAdd real-world generative reach, acquisition and CRM/order outcomesDemand/commercial bridge
Causal reviewCompare treatment with appropriate counterfactualTreatment-effect conclusion
Next experimentRetain learning and identify the next diagnosed deficiencyRepeatable improvement cycle

The detailed research expressed this as a twelve-week programme.

The exact calendar is less important than the sequence:

Instrument first. Establish the baseline and normal variation. Intervene deliberately. Verify source and answer movement. Test robustness. Then connect that movement to downstream commercial outcomes.

Methodological output

What did this work change in Kojable's measurement model?

The programme began with a useful measurement question:

Did the company appear, and which sources were cited?

The research showed why a mature answer-alignment programme needs to separate more stages.

It now distinguishes:

presence from source uptake;

source uptake from information reflection;

information reflection from accuracy and completeness;

visibility from appropriate recommendation;

one-period movement from persistent movement;

one prompt or system from robust cross-system behaviour;

and answer-layer improvement from commercial incrementality.

That separation is the principal methodological output of the case study.

What does this case study support?

The study directly supports several conclusions.

Kojable developed and applied a repeatable AI-answer measurement methodology to two anonymised B2B programmes.

The programmes used:

  • fixed buyer-question panels;
  • five AI/search surfaces;
  • three repetitions per question/surface combination;
  • US-English conditions;
  • retained answer and citation records;
  • explicit separation of branded and unbranded questions;
  • explicit handling of technical failures.

Company A appeared in 4 of 164 usable unprompted answers.

Company B appeared in 56 of 105, with some tracked intent cells reaching 15 of 15 mentions.

Those findings demonstrate that two valid visibility measurements can lead to different next questions.

The work also produced a wider methodology for prospective interventions.

It did not retrospectively create historical recommendation-rate or claim-level accuracy percentages that were absent from the original baseline.

It did not establish that a particular content intervention caused later movement.

It did not attribute incremental pipeline or revenue to AI visibility.

Those are later measurement layers with their own definitions and evidence requirements.

Next action

What should a team do with an AI visibility baseline?

A Company A-type state should trigger diagnosis before optimisation.

Low unprompted presence tells the team that the company rarely enters relevant measured answers.

It does not reveal the hidden mechanism automatically.

The next task is to examine:

  • how clearly the company is represented in public evidence;
  • which source patterns recur;
  • whether relevant buyer-intent evidence exists;
  • whether entity/category information is clear;
  • what intervention can be tested without assuming the answer in advance.

A Company B high-presence state requires a different response.

Where the company already appears consistently, maximising raw mention rate offers little additional diagnostic value.

The next measurement should shift towards:

  • accuracy;
  • completeness;
  • differentiation;
  • source support;
  • appropriate recommendation;
  • persistence;
  • robustness.

After an intervention, the programme should verify whether:

  1. the intended page surfaced;
  2. the changed information entered the answer;
  3. answer quality remained acceptable;
  4. recommendations remained appropriate;
  5. the result persisted.

That is the practical connection to the Kojable operating model:

Monitor → Diagnose → Improve → Verify.

FAQ

Frequently asked questions

What is AI search visibility measurement?

AI search visibility measurement is the systematic observation of whether and how a company appears in relevant AI-generated buyer answers.

A useful visibility metric needs a defined question set, eligible population, numerator, denominator, AI/search surfaces, repetitions, market, language, collection period and failure-handling rule.

The resulting visibility percentage establishes an observed state. It does not by itself diagnose why that state exists.

Should branded prompts count towards AI visibility?

They should be separated when the objective is unprompted buyer discovery.

A prompt that explicitly names the company measures how it is represented once it is already known.

An unbranded buyer question measures whether the company enters the answer without the user supplying its name.

Both are useful, but they answer different questions.

How many times should an AI prompt be tested?

This study used three repetitions per question/surface combination.

Three is not a universal standard.

The appropriate number depends on the research objective, expected variability, number of independent buyer intents, measurement cost and desired statistical precision.

Repeated evaluation matters because generative outputs vary even when the input remains the same. See OpenAI's evaluation best practices.

Are AI mentions, citations and recommendations the same thing?

No.

A mention means the company appears in the visible answer.

A citation means a source is visibly attributed.

A recommendation requires a separate judgement that the answer is presenting the company as an appropriate option for the buyer's stated need.

They should be measured separately.

Can AI visibility be connected to pipeline or revenue?

Yes, but answer-layer movement and commercial incrementality should remain separate measurements.

A programme can connect controlled answer benchmarks with generative-search exposure, AI-assistant traffic, qualified actions, opportunities, pipeline and revenue.

Establishing incremental commercial value still requires an appropriate lag window and counterfactual design.

Provenance

Sources and methodology

Primary evidence

  • Kojable anonymised AI-answer baseline records, 10–11 September 2026.

    These retained machine-recorded observations contain the prompt, answer, citation, AI/search surface, repetition and collection metadata used for the baseline findings.

  • Kojable GEO / Answer Search Optimisation Measurement Framework.

    This is the methodological source for the staged measurement system, prospective metrics, evidence hierarchy, reporting cadence and intervention decision rules presented in this case study.

External methodological and platform context

Provenance note: The case-study findings come from Kojable's retained machine-recorded baseline observations. External sources provide methodological and platform context; they are not the source of the Company A or Company B results.

Conclusion

Conclusion

The most important result from these two B2B programmes was not that one visibility percentage was larger than another.

It was that a visibility score stopped being useful as soon as it was treated as the diagnosis.

Company A's 2.4% result first raised an answer-level buyer-discovery question.

Company B's 53.3% overall result concealed individual buyer intents where raw mention presence had already reached 15 of 15.

Those states require different next measurements.

The resulting framework therefore asks a sequence of progressively more useful questions:

Can the company be discovered?

Does the intended source surface?

Does the changed information enter the answer?

Is the answer accurate and complete?

Is the recommendation appropriate?

Does the effect persist?

Does it survive prompt and system variation?

Does it ultimately produce measurable buyer and commercial value?

Each question needs its own evidence.

That makes AI search visibility useful for what it should be:

a starting point for diagnosis, not the end of the measurement system.

For B2B companies, the operating loop is:

Monitor the answer. Diagnose the gap. Improve the relevant evidence. Verify what changed.

Run your free AI brand audit

Establish your current baseline across the buyer questions that matter, then identify which answer-alignment problem deserves attention first.

Author

About Kojable

KojableAI answer alignment platform · Case Studies