Measurement case study · AI search visibility
Same Visibility Metric, Different Problem: What Two B2B Baselines Revealed
AI search visibility is useful as a baseline, not a diagnosis. In two anonymised US-English B2B benchmarks, Company A appeared in 4 of 164 usable unprompted answers, while Company B appeared in 56 of 105. The panels were different and are not a head-to-head comparison. Their value was diagnostic: Company A showed low answer-level presence across its tracked buyer questions, while some Company B intents were already at 15 of 15 mentions. Those different states led Kojable to develop a staged measurement framework spanning source uptake, information reflection, answer quality, robustness and commercial validation.
Case question: How should a B2B company interpret an AI search visibility baseline, and what should it measure next?
- Programmes
- 2 anonymised B2B benchmarks
- Planned checks
- 330
- Surfaces
- 5 AI/search surfaces
- Repetitions
- 3 per question/surface
Measurement canvas
From buyer questions to the next measurement decision
The important finding was not that one company had a bigger percentage than the other. It was that the same type of visibility metric exposed two different measurement problems.
- Benchmarks
- 2
- Planned checks
- 330
- Surfaces
- 5
- Repetitions
- 3
-
01
Buyer questions
The benchmark used deliberately selected buyer questions rather than attempting to estimate every prompt a buyer might ask.
-
02
Repeated observations
Each question was measured three times across five user-facing AI/search surfaces.
-
03
Unprompted discovery + citations
Record whether the company appeared and which sources were visibly cited.
-
04
Diagnosis + next measurement
The next metric should depend on the diagnosed problem.
The baseline describes the observed answer-layer starting state. It does not by itself diagnose the hidden upstream cause or establish intervention causality.
Programme context
Kojable developed the measurement methodology and conducted the AI-answer and citation analysis as the research and intelligence layer of an anonymised client programme. The methodology covers prompt design, source discovery, extraction, scoring, orchestration, validation and quality control.
The underlying measurement records retain the exact questions, answer text, citations, AI/search surface, repetition number and machine-recorded timestamps.
The important finding was not that one company had a bigger percentage than the other. It was that the same type of visibility metric exposed two different measurement problems.
That changed the question from:
How visible is the company?
to:
What does this starting state tell us to investigate next, how should that change be measured, and what evidence would justify a stronger conclusion?
Baseline design
How was AI search visibility measured?
The benchmark used deliberately selected buyer questions rather than attempting to estimate every prompt a buyer might ask.
Each question was measured three times across five user-facing AI/search surfaces:
- ChatGPT
- Gemini
- Google AI Overview
- Google AI Mode
- Claude
The study was configured for the United States market and English-language responses.
| Measurement field | Company A | Company B |
|---|---|---|
| Buyer questions | 12 | 10 |
| AI/search surfaces | 5 | 5 |
| Repetitions per question/surface | 3 | 3 |
| Planned observations | 180 | 150 |
| Collection date | 10 Sep 2026 | 11 Sep 2026 |
| Market | United States | United States |
| Language | English | English |
| Headline discovery metric | Unprompted Brand Discovery Rate | Unprompted Brand Discovery Rate |
Each repetition remained a separate observation rather than being averaged into a single generated answer.
That matters because generative outputs vary. OpenAI's current evaluation guidance explicitly notes that models can produce different outputs from the same input, recommends task-specific and continuous evaluation, and advises combining numerical metrics with human judgement.
Three repetitions were the design used in this benchmark. They should not be treated as a universal optimum for every AI visibility study.
Measurement contract
For this case study, the visibility percentages should be interpreted only with their measurement contract:
questions → eligibility → numerator → denominator → surfaces → repetitions → market/language → collection window → failure handling
Without those elements, a percentage labelled “AI visibility” can represent materially different things.
A technical collection failure was excluded from answer-level outcome metrics. A valid observation in which an AI Overview was not shown was retained as a valid surface outcome and treated differently from a collection failure.
What is Unprompted Brand Discovery Rate?
Unprompted Brand Discovery Rate is the share of usable answers to eligible buyer questions that do not name the company in which the company nevertheless appears.
Formula
eligible unbranded answers naming the company ÷ usable eligible unbranded answers
The analytical unit is one recorded answer observation.
Questions explicitly naming the client remain useful for analysing known-brand representation, but they are excluded from the unprompted buyer-discovery denominator.
That distinction prevents visibility from being inflated simply because the user supplied the company name.
A question asking:
“What does Company X offer?”
answers a different question from:
“Which providers can solve this problem?”
The first tests representation once the company is already known. The second tests whether the company enters the buyer's answer without being supplied by the user.
Related Kojable research, Branded AI Visibility Is Not Market Visibility, examines the same denominator problem from a different dataset.
Company A baseline
What did Company A's baseline show?
Company A's panel contained 11 unprompted buyer questions after one explicitly branded question was separated from the discovery metric.
That produced:
11 questions × 5 surfaces × 3 repetitions = 165 planned unprompted observations
One Google AI Overview collection failed technically.
It was excluded rather than scored as a negative brand outcome, leaving 164 usable unprompted observations.
Company A appeared in 4 of those 164 answers, or 2.4%.
| Surface | Brand mentions | Usable unprompted answers | Discovery rate |
|---|---|---|---|
| ChatGPT | 0 | 33 | 0.0% |
| Gemini | 0 | 33 | 0.0% |
| Google AI Overview | 1 | 32 | 3.1% |
| Google AI Mode | 1 | 33 | 3.0% |
| Claude | 2 | 33 | 6.1% |
| Total | 4 | 164 | 2.4% |
Figure 1 — Company A: unprompted brand discovery by AI/search surface
Figure annotation: 4 of 164 usable answers · 2.4%
Source: Kojable anonymised AI-answer measurement programme, 10 September 2026.
The defensible finding is:
Company A had low unprompted answer-level presence across its selected buyer-question benchmark.
It does not tell us why.
A company can be absent from an answer because of discoverability, retrieval, selection, synthesis, the available evidence environment or another factor that cannot be established from the endpoint alone.
The result identifies where the problem is visible at the answer layer; it does not establish whether the upstream constraint is discoverability, retrieval, selection, synthesis or the public evidence environment.
So the baseline identifies where the problem is visible, not automatically what caused it.
Company B baseline
What did Company B's baseline show?
Company B contained seven unprompted buyer questions after three explicitly brand-named questions were separated.
That produced:
7 questions × 5 surfaces × 3 repetitions = 105 unprompted observations
Company B appeared in 56 of 105 answers, or 53.3%.
| Surface | Brand mentions | Unprompted answers | Discovery rate |
|---|---|---|---|
| ChatGPT | 7 | 21 | 33.3% |
| Gemini | 13 | 21 | 61.9% |
| Google AI Overview | 14 | 21 | 66.7% |
| Google AI Mode | 12 | 21 | 57.1% |
| Claude | 10 | 21 | 47.6% |
| Total | 56 | 105 | 53.3% |
Figure 2 — Company B: unprompted brand discovery by AI/search surface
Figure annotation: 56 of 105 answers · 53.3%
Source: Kojable anonymised AI-answer measurement programme, 11 September 2026.
The more useful finding came from looking below the overall percentage.
Two Company B buyer-intent areas produced mentions in 15 of 15 measured answers. Another tracked intent produced none.
That changes the measurement problem.
If a company already appears in 15 of 15 answers for an intent, raw mention visibility has no remaining headroom inside that benchmark.
The next question should not be:
How do we increase the mention rate?
It should become:
Is the company represented accurately, differentiated appropriately, supported by useful evidence and recommended where it genuinely fits?
Interpretation
Why didn't 2.4% and 53.3% mean the same thing?
Because a visibility percentage describes an observed state.
It does not diagnose that state.
| Observed state | What the result tells us | What it does not establish | More useful next question |
|---|---|---|---|
| Low unprompted presence | The company rarely enters the measured buyer answers | Why the company is absent | Which evidence, entity or information gap needs investigation? |
| High or ceiling-level presence | Brand inclusion is already strong for that intent | That representation quality is good | Is the company described accurately and distinctively? |
| Citation without visible brand | A source was attributed | That the brand became visible | What did the source actually support? |
| Brand mention without target citation | The company entered the answer | That the intended page was used | Which evidence paths are associated with the answer? |
| Strong presence with unknown recommendation quality | The brand is visible | That it is an appropriate recommendation | Does the company genuinely fit the buyer's requirements? |
| One post-change increase | The observed result moved | That the intervention caused it | Does movement persist against a credible comparison? |
This is the central finding of the case study:
Answer
A visibility metric establishes the starting state. The next metric should depend on the diagnosed problem.
Kojable is an AI answer alignment platform for B2B companies. Visibility is one measurement dimension inside the wider operating model:
Monitor → Diagnose → Improve → Verify.
Measurement gap
Why were visibility and citation counts not enough?
The original programme could answer several useful questions:
- Did the company appear?
- On which AI/search surface?
- For which buyer question?
- Which sources were visibly cited?
But analysing those results exposed questions those metrics could not answer on their own.
A brand-related citation did not necessarily mean the intended intervention page appeared.
A cited page did not necessarily mean the important information on that page entered the generated answer.
A mention did not necessarily constitute a recommendation.
A recommendation would not necessarily be useful if the company did not objectively fit the buyer's requirements.
And an increase in AI visibility would not establish commercial incrementality.
The measurement system therefore needed to evolve from:
Did the company appear and how many citations were returned?
towards:
Did the intended source surface, did the changed information enter the answer, did representation quality remain acceptable, did the result persist and did any downstream buyer behaviour change?
Expanded framework
The measurement system that emerged from the baseline work
Rather than turning one visibility score into a larger composite score, Kojable separated the programme into distinct measurement questions.
| Layer | Metric | Question answered |
|---|---|---|
| Buyer discovery | Unprompted Brand Discovery Rate | Does the company enter relevant buyer answers without being named? |
| Source uptake | Target-Source Uptake Rate | Does the exact intervention URL begin appearing? |
| Information reflection | Changed-Information Reflection Rate | Does the introduced or corrected information actually enter the answer? |
| Evidence quality | Citation Support Rate | Does the cited source genuinely support the associated claim? |
| Representation quality | Answer Accuracy | Are material factual claims correct? |
| Answer Completeness | Are the required decision-relevant facts included? | |
| Recommendation quality | Appropriate Recommendation Rate | Is the company recommended where it objectively fits? |
| Inappropriate / Misrepresentation Rate | Is optimisation producing unsuitable or inaccurate recommendations? | |
| Recommendation Share of Voice | How frequently is the company recommended relative to named alternatives? | |
| Durability | Persistence | Does the effect hold across later periods? |
| Prompt Robustness | Does it survive semantically equivalent prompt variants? | |
| Cross-System Robustness | Does the result hold across more than one AI/search surface? | |
| System-Transition Retention | Does the effect survive a material model/product transition? | |
| Latency to Uptake | How long does sustained source or information uptake take? | |
| Observed market reach | Generative-AI observed reach | Is the site appearing in real-world Google generative-search exposure? |
| Qualified acquisition | AI Assistant qualified value | Are AI-assistant visits producing useful buyer actions? |
| Commercial validation | Pipeline/revenue/contribution treatment effect | Did an intervention create incremental business value? |
Only the baseline discovery and original citation environment were historically measured in these two programmes.
The later measures are the prospective framework developed from the analysis, not retrospective results.
-
Historically measured
Measured baseline
Unprompted buyer discovery
Baseline citation environment
-
Prospective
Source and answer measurement
Target-source uptake → Changed-information reflection → Citation support → Accuracy / completeness → Appropriate recommendation
-
Prospective
Durability
Persistence → Prompt robustness → Cross-system robustness → System-transition retention
-
Prospective
Downstream validation
Observed reach → Qualified buyer action → Commercial validation
Only Unprompted Brand Discovery and the original citation environment were measured historically in these baselines. Later stages are the prospective measurement framework developed from the analysis.
This is a measurement architecture. It does not imply that every commercial AI product exposes each hidden processing stage.
What is Target-Source Uptake Rate?
General citation volume is useful for understanding the source environment.
It becomes less specific when the research question concerns a particular intervention.
Target-Source Uptake Rate is the share of eligible answers that cite the exact URL affected by an intervention.
Formula
eligible answers citing the exact intervention URL ÷ eligible answers
Suppose a team publishes a new research page to clarify an outdated product fact.
If another company page starts appearing more frequently, total first-party citation activity may rise without showing whether the intended research page has begun surfacing.
Target-source measurement keeps the outcome connected to the specific asset being tested.
It should therefore be registered alongside:
- intervention URL;
- deployment date;
- affected buyer questions;
- changed claim or evidence;
- crawl/index confirmation where applicable;
- post-intervention observation window.
Target-Source Uptake Rate was not imposed retrospectively on these baselines. It is a prospective intervention metric.
What is Changed-Information Reflection Rate?
A target page being cited still does not answer the most important information question:
Did the changed information enter the answer?
Kojable therefore defines Changed-Information Reflection Rate as:
The share of eligible answers that correctly express a pre-specified fact introduced or corrected through an intervention.
Formula
eligible answers correctly expressing the changed fact ÷ eligible answers
The target fact should be defined before deployment.
For example, if a page contains an ambiguous capability statement and the intervention adds a verified clarification, measurement should not stop at whether the page appears as a citation.
It should also ask:
Does the subsequent answer correctly reflect the clarification?
This lets the programme distinguish two materially different results.
Target citation rises, reflection does not
The page may be surfacing, but the intended information is not consistently appearing in answers.
That suggests a need to investigate relevance, clarity, source context or other information pathways.
Reflection rises, target citation does not
The changed information may be reaching answers through another public source or without visible attribution to the intended page.
That warrants investigation rather than a claim of direct citation causality.
Neither outcome alone proves the hidden causal path, but the distinction provides substantially more diagnostic value than citation volume alone.
Why use a more specific citation taxonomy?
A generic “brand-related citation” flag becomes insufficient once the programme starts measuring interventions.
Different source relationships mean different things.
Kojable's expanded framework separates:
| Citation class | Meaning |
|---|---|
| Exact target URL | The specific page affected by the intervention |
| First-party brand citation | A citation on the company's owned domain |
| Explicit third-party brand citation | A third-party title, source or URL explicitly identifying the company |
| Content-verified brand citation | A third-party page that materially discusses the company even when the title or URL does not |
| Competitor/comparison citation | A comparison or competitor source discussing the company |
| Supporting non-brand source | A source supporting the category or answer without materially discussing the company |
This avoids treating all citations as interchangeable.
It also avoids several common inference errors.
A third-party comparison citation is not the same event as an exact first-party target page surfacing.
A supporting category source does not necessarily create company visibility.
And a visible citation does not, by itself, prove that the source caused the answer.
Historical series should not be silently rewritten when citation definitions change. Where classifications are improved, retained evidence can be rescored explicitly or old and new definitions can be reported as separate series.
Are mentions, citations and recommendations the same metric?
No.
Mention ≠ citation ≠ recommendation.
A mention records that the company appeared in the visible answer.
A citation records that the surface visibly attributed a source.
A recommendation requires a separate judgement that the answer is presenting the company as an appropriate option for the buyer's stated need.
The original baseline measured mentions and citation behaviour.
It did not contain a separate historical recommendation rubric, so this case study does not retrospectively create a recommendation percentage.
That is important because an optimisation programme should not reward a system simply for recommending the company more often.
A stronger future measure needs an eligibility rule.
For example, the programme might define eligibility using verified:
- use cases;
- product capabilities;
- company size;
- deployment requirements;
- integrations;
- security requirements;
- pricing constraints.
The result can then be classified as an appropriate recommendation only where the company genuinely fits.
How should answer quality be protected?
More visibility is not automatically better answer alignment.
A company can become more visible while being represented incorrectly.
A recommendation can increase while becoming less appropriate.
An intervention can improve a visibility proxy while reducing answer quality.
The framework therefore introduces quality guardrails.
Citation Support Rate
Checked cited claims genuinely supported by their cited source ÷ cited claims checked
This asks whether the visible source actually supports the associated statement.
It does not claim to reveal the system's hidden reasoning.
Answer Accuracy
Correct material factual claims ÷ material claims scored
The material claims and scoring standard must be defined before evaluation.
Answer Completeness
Required decision-relevant facts correctly covered ÷ pre-registered required facts
This catches omissions as well as outright factual errors.
Appropriate Recommendation Rate
Eligible buyer answers in which the company is appropriately recommended ÷ eligible recommendation-intent answers
The eligibility rule must be fixed before collection.
Inappropriate Recommendation / Misrepresentation Rate
This is the guardrail.
If recommendation visibility increases while factual accuracy or buyer fit decreases, that is harmful optimisation, not success.
Durability
How do you know whether an improvement is durable?
A movement appearing once on one surface for one wording is a weaker result than one that survives time, prompt variation and system changes.
The framework therefore separates the size of an initial movement from its durability.
| Measure | Question answered |
|---|---|
| Persistence | Does the result continue across later measurement periods? |
| Prompt Robustness | Does it survive semantically equivalent buyer questions? |
| Cross-System Robustness | Does the effect appear in the intended direction across more than one measured surface? |
| System-Transition Retention | What happens when a model or product materially changes? |
| Latency to Uptake | How long does it take from deployment or index confirmation to sustained source/information uptake? |
Instead of attempting to predict exactly what an undisclosed future model will prefer, the programme can ask:
Which interventions continue to produce useful, accurate effects across prompts, systems, time and model/product transitions?
That is measurable resilience rather than speculative future-model optimisation.
Evidence strength
When can an AI visibility change support a causal claim?
A post-intervention increase is an observation.
Timing alone does not establish cause.
Kojable therefore separates evidence states.
Observed
A post-change association appears.
Supported
The exact source and/or changed information repeatedly appears.
Controlled
A credible treatment/comparison design shows differential movement.
Causal
A randomised or sufficiently strong quasi-experimental design supports a treatment-effect claim.
Commercially causal
The downstream business outcome also survives an appropriate counterfactual analysis.
The two historical baseline programmes did not progress through or achieve every evidence level.
This helps match claim strength to evidence strength.
A target URL beginning to appear after deployment is more informative than total citation volume.
A target URL appearing alongside correct reflection of a newly introduced fact is stronger again.
A credible treatment/control design supports stronger intervention inference.
Commercial incrementality needs its own downstream comparison.
The framework does not require teams to stay silent until a perfect experiment exists. It requires them to describe the evidence they actually have.
What should different post-intervention patterns mean?
The framework turns combinations of measures into decisions.
| Observed pattern | Interpretation | Next action |
|---|---|---|
| Target citation ↑, reflection unchanged | The source is surfacing, but the intended information is not consistently entering answers | Examine content clarity, relevance and citation context |
| Target citation unchanged, reflection ↑ | Information may be entering through another evidence path | Investigate third-party propagation and other sources |
| Citation ↑, reflection ↑, appropriate recommendation ↑ and persists | Strong answer-layer evidence | Replicate on another topic cluster and test commercial linkage |
| AI-answer metrics ↑, commercial outcomes unchanged | Answer-layer movement is established; business value remains unproven | Continue the commercial lag window and inspect intent/conversion quality |
| AI-answer and commercial treatment effects ↑ against credible controls | Strongest evidence state | Consider scaling while retaining holdouts or controls |
| Single-period jump | Insufficient durability evidence | Repeat before drawing a strategic conclusion |
| Recommendation ↑, accuracy or fit ↓ | Harmful optimisation | Correct or roll back the intervention |
These are future decision rules, not retrospective claims about Company A or Company B.
Reporting cadence
How should reporting change as the programme matures?
Not every metric matures on the same timescale.
A weekly AI-answer report should not pretend to prove quarterly revenue incrementality.
A quarterly business review should not be dominated by screenshots from individual prompts.
The framework therefore separates reporting by decision horizon.
| Cadence | Primary job | Example measures |
|---|---|---|
| Weekly | Detect answer-level movement and operational issues | Discovery, target-source uptake, reflection, accuracy, citation support, recommendation fit, unavailable runs, model/system anomalies |
| Monthly | Connect answer movement with observed reach and qualified demand | Google generative-AI impressions, branded/non-branded demand, Organic Search, AI Assistant traffic, qualified visits, trials/demos/leads |
| Quarterly | Evaluate commercial and causal evidence | Treatment/control effects, opportunities, pipeline, new ARR/bookings, revenue, contribution, cohort outcomes |
Weekly
Discovery, target-source uptake, reflection, accuracy, citation support, recommendation fit, unavailable runs, system/model anomalies.
Monthly
Generative-AI impressions, branded/non-branded demand, Organic Search, AI Assistant traffic, qualified visits, trials/demos/leads.
Quarterly
Treatment/control effects, opportunities, pipeline, new ARR/bookings, revenue, contribution, cohort outcomes.
Answer-layer movement, observed reach and commercial incrementality remain separate reporting layers.
Google's Search Console now provides dedicated generative-AI performance reporting, including impressions plus page, country, device and date dimensions. Google says the reports had rolled out worldwide by 31 August 2026.
That observational Google layer should complement, not replace, the controlled buyer-question benchmark.
The two populations are different.
The benchmark exposes the actual answer, brand presence and citations for a fixed question panel.
Search Console shows how URLs from the site are appearing within Google's generative Search features in real usage.
GA4 now also defines an AI Assistant channel for sources including ChatGPT and Gemini. Google's AI Overviews and AI Mode are excluded from that channel and included within Organic Search. See Google's Default channel group documentation.
That enables a future measurement chain such as:
controlled answer benchmark → observed generative reach → AI/organic acquisition → qualified action → pipeline/revenue
The stages should still be reported separately.
Why does the measurement contract matter?
AI visibility is becoming an industry measurement category, but methodologies still differ substantially.
IAB's August 2026 Measuring Visibility in the AI Era guidance was developed in response to that problem. IAB says more than 20 companies now sell AI visibility measurement tools and that differing methodologies can produce different answers for the same brand or publisher.
That supports an important conclusion from the Kojable programmes:
An AI visibility percentage has limited meaning without its measurement contract.
At minimum, the reader needs to know:
- Question panel
which questions were measured;
- Eligibility
which answers were eligible;
- Numerator
the numerator;
- Denominator
the denominator;
- Surfaces
which AI/search surfaces were included;
- Repetitions
how many repetitions were run;
- Market and language
market and language;
- Collection dates
collection dates;
- Failure handling
how unavailable observations were handled.
That is also why Company A's 2.4% and Company B's 53.3% cannot be turned into a direct brand ranking.
Both percentages can be correctly calculated while answering different benchmark questions.
Operating sequence
How does the measurement programme evolve after the baseline?
The analysis produced not only new metrics but an operating sequence.
| Phase | Purpose | Output |
|---|---|---|
| Foundation | Freeze prompts, preserve raw responses and citations, standardise definitions | Reproducible baseline |
| Calibration | Estimate ordinary variation and establish scoring rubrics | Variance and scoring model |
| Experiment | Register changed claims, intervention URLs and comparisons | Testable intervention hypothesis |
| Post-intervention measurement | Measure target-source uptake and changed-information reflection | Answer-layer evidence |
| Robustness | Test persistence, paraphrases and cross-system behaviour | Durability evidence |
| Commercial linkage | Add real-world generative reach, acquisition and CRM/order outcomes | Demand/commercial bridge |
| Causal review | Compare treatment with appropriate counterfactual | Treatment-effect conclusion |
| Next experiment | Retain learning and identify the next diagnosed deficiency | Repeatable improvement cycle |
The detailed research expressed this as a twelve-week programme.
The exact calendar is less important than the sequence:
Instrument first. Establish the baseline and normal variation. Intervene deliberately. Verify source and answer movement. Test robustness. Then connect that movement to downstream commercial outcomes.
Methodological output
What did this work change in Kojable's measurement model?
The programme began with a useful measurement question:
Did the company appear, and which sources were cited?
The research showed why a mature answer-alignment programme needs to separate more stages.
It now distinguishes:
presence from source uptake;
source uptake from information reflection;
information reflection from accuracy and completeness;
visibility from appropriate recommendation;
one-period movement from persistent movement;
one prompt or system from robust cross-system behaviour;
and answer-layer improvement from commercial incrementality.
That separation is the principal methodological output of the case study.
What does this case study support?
The study directly supports several conclusions.
Kojable developed and applied a repeatable AI-answer measurement methodology to two anonymised B2B programmes.
The programmes used:
- fixed buyer-question panels;
- five AI/search surfaces;
- three repetitions per question/surface combination;
- US-English conditions;
- retained answer and citation records;
- explicit separation of branded and unbranded questions;
- explicit handling of technical failures.
Company A appeared in 4 of 164 usable unprompted answers.
Company B appeared in 56 of 105, with some tracked intent cells reaching 15 of 15 mentions.
Those findings demonstrate that two valid visibility measurements can lead to different next questions.
The work also produced a wider methodology for prospective interventions.
It did not retrospectively create historical recommendation-rate or claim-level accuracy percentages that were absent from the original baseline.
It did not establish that a particular content intervention caused later movement.
It did not attribute incremental pipeline or revenue to AI visibility.
Those are later measurement layers with their own definitions and evidence requirements.
Next action
What should a team do with an AI visibility baseline?
A Company A-type state should trigger diagnosis before optimisation.
Low unprompted presence tells the team that the company rarely enters relevant measured answers.
It does not reveal the hidden mechanism automatically.
The next task is to examine:
- how clearly the company is represented in public evidence;
- which source patterns recur;
- whether relevant buyer-intent evidence exists;
- whether entity/category information is clear;
- what intervention can be tested without assuming the answer in advance.
A Company B high-presence state requires a different response.
Where the company already appears consistently, maximising raw mention rate offers little additional diagnostic value.
The next measurement should shift towards:
- accuracy;
- completeness;
- differentiation;
- source support;
- appropriate recommendation;
- persistence;
- robustness.
After an intervention, the programme should verify whether:
- the intended page surfaced;
- the changed information entered the answer;
- answer quality remained acceptable;
- recommendations remained appropriate;
- the result persisted.
That is the practical connection to the Kojable operating model:
Monitor → Diagnose → Improve → Verify.
FAQ
Frequently asked questions
What is AI search visibility measurement?
AI search visibility measurement is the systematic observation of whether and how a company appears in relevant AI-generated buyer answers.
A useful visibility metric needs a defined question set, eligible population, numerator, denominator, AI/search surfaces, repetitions, market, language, collection period and failure-handling rule.
The resulting visibility percentage establishes an observed state. It does not by itself diagnose why that state exists.
Should branded prompts count towards AI visibility?
They should be separated when the objective is unprompted buyer discovery.
A prompt that explicitly names the company measures how it is represented once it is already known.
An unbranded buyer question measures whether the company enters the answer without the user supplying its name.
Both are useful, but they answer different questions.
How many times should an AI prompt be tested?
This study used three repetitions per question/surface combination.
Three is not a universal standard.
The appropriate number depends on the research objective, expected variability, number of independent buyer intents, measurement cost and desired statistical precision.
Repeated evaluation matters because generative outputs vary even when the input remains the same. See OpenAI's evaluation best practices.
Are AI mentions, citations and recommendations the same thing?
No.
A mention means the company appears in the visible answer.
A citation means a source is visibly attributed.
A recommendation requires a separate judgement that the answer is presenting the company as an appropriate option for the buyer's stated need.
They should be measured separately.
Can AI visibility be connected to pipeline or revenue?
Yes, but answer-layer movement and commercial incrementality should remain separate measurements.
A programme can connect controlled answer benchmarks with generative-search exposure, AI-assistant traffic, qualified actions, opportunities, pipeline and revenue.
Establishing incremental commercial value still requires an appropriate lag window and counterfactual design.
Provenance
Sources and methodology
Primary evidence
- Kojable anonymised AI-answer baseline records, 10–11 September 2026.
These retained machine-recorded observations contain the prompt, answer, citation, AI/search surface, repetition and collection metadata used for the baseline findings.
- Kojable GEO / Answer Search Optimisation Measurement Framework.
This is the methodological source for the staged measurement system, prospective metrics, evidence hierarchy, reporting cadence and intervention decision rules presented in this case study.
External methodological and platform context
- OpenAI, Evaluation best practices
- Google Search Central, Introducing Search Generative AI performance reports in Search Console
- Google Analytics, Default channel group
- IAB, Measuring Visibility in the AI Era
Provenance note: The case-study findings come from Kojable's retained machine-recorded baseline observations. External sources provide methodological and platform context; they are not the source of the Company A or Company B results.
Conclusion
Conclusion
The most important result from these two B2B programmes was not that one visibility percentage was larger than another.
It was that a visibility score stopped being useful as soon as it was treated as the diagnosis.
Company A's 2.4% result first raised an answer-level buyer-discovery question.
Company B's 53.3% overall result concealed individual buyer intents where raw mention presence had already reached 15 of 15.
Those states require different next measurements.
The resulting framework therefore asks a sequence of progressively more useful questions:
Can the company be discovered?
Does the intended source surface?
Does the changed information enter the answer?
Is the answer accurate and complete?
Is the recommendation appropriate?
Does the effect persist?
Does it survive prompt and system variation?
Does it ultimately produce measurable buyer and commercial value?
Each question needs its own evidence.
That makes AI search visibility useful for what it should be:
a starting point for diagnosis, not the end of the measurement system.
For B2B companies, the operating loop is:
Monitor the answer. Diagnose the gap. Improve the relevant evidence. Verify what changed.
Run your free AI brand audit
Establish your current baseline across the buyer questions that matter, then identify which answer-alignment problem deserves attention first.