Kojable Blog reference entry
Topic-Level AI Visibility: How to Compare Topics Without Creating False Gaps
Topic-level AI visibility measures a company's presence or prominence within a defined subject or buyer-question cluster under an explicit prompt set, platform or surface, eligibility rule, run design and time window.
Category AI Search Guides
Also known as AI visibility by topic, topic visibility measurement, AI topic visibility
Quick answer: Do not assume two topic-level AI visibility results are directly comparable. A weaker topic may reflect a real presence gap, but it may also reflect differences in prompt mix, platform, retrieval context, relevance rules, sample breadth or collection status. Compare topics only after checking that the metric, buyer-question scope, platforms, eligibility rules, observation window and denominators are sufficiently aligned. If they are not, report the result as directional, qualified or non-comparable rather than calling it a visibility gap. Diagnosis should begin only after a comparable difference has been established.
A company can look strong on one AI topic and weak on another without the difference necessarily representing a genuine topic-level visibility problem.
The reason is simple: a percentage inherits the design of the observations behind it.
If one topic contains broad informational questions while another contains evidence-heavy vendor comparisons, their results may reflect different information needs. If one topic is heavily represented on one AI platform and another has a different platform mix, a pooled comparison can conceal those platform differences. If one result is built from many distinct buyer questions and another comes from repeated versions of a narrow prompt, the two percentages do not have the same breadth.
AI visibility should therefore be interpreted through the measurement population behind the score, not only the displayed percentage.
Kojable Research provides a useful example. In its study of AI citations by buyer-question type, a large raw difference on Gemini became much smaller when the comparison was restricted to a more comparable retrieval-need cohort. The study does not show that every topic comparison behaves this way. It shows why raw category differences should be tested for comparability before they are interpreted as independent gaps.
This study measured visible citation exposure rather than company visibility overall. It is used here as an example of why raw segment differences can change when the underlying comparison populations are made more comparable.
Read the canonical Kojable Research: AI Citations by Query Type
That creates a measurement question teams should answer before diagnosis:
Are the two topic results measuring sufficiently similar populations to justify saying one topic is weaker?
01 · Definition
What is topic-level AI visibility?
Topic-level AI visibility measures a company's presence or prominence within a defined subject or buyer-question cluster under an explicit measurement design.
The important phrase is under an explicit measurement design.
A topic score is not just a label and a percentage. It is the result of choices about:
- which buyer questions belong to the topic;
- which prompts represent those questions;
- which AI platforms or surfaces are included;
- what counts as a relevant response;
- what event counts as visibility;
- how many observations exist;
- how broad the prompt set is;
- when the observations were collected;
- and which responses are eligible for the denominator.
That means two values called “topic visibility” can look comparable while actually representing different analytical populations.
For example, one topic might contain questions about vendor evaluation and implementation requirements while another contains primarily educational questions. Those prompt groups can have different information needs and different observable citation behaviour. Kojable's query-type research found that buyer-question type was informative in some platform contexts, but the relationship was not universal and overlapped substantially with retrieval need.
A topic score is therefore useful only when the reader can understand what sits underneath it.
02 · Comparison contract
When are two topic-level AI visibility results actually comparable?
Two topic results are sufficiently comparable when the metric definition and the populations behind the measurements are aligned closely enough for the difference to answer the intended question.
For this practical application, use a Topic Comparison Contract before calling one topic stronger or weaker. It is a reporting framework, not a validated statistical test.
| Comparison dimension | More comparable when | Risk of an apparent or inflated gap when |
|---|---|---|
| Metric | Both topics measure the same defined outcome | One result reflects mentions while another is effectively interpreted as citations or recommendations |
| Buyer-question scope | The topic groups represent clearly defined and intentionally comparable analytical units | One topic contains broad informational questions while another contains evidence-heavy comparison questions |
| Platform or surface | Results are compared within platform or the platform mix is explicitly controlled | Different platform mixes are pooled into one topic comparison |
| Retrieval context | Material differences in information need are considered or stratified | One topic systematically requires more current, comparative or technical evidence |
| Prompt breadth | Both topics contain meaningful coverage of distinct buyer questions | One topic contains diverse questions while another repeats a narrow prompt design |
| Eligibility rules | The same relevance rules determine which observations enter the calculation | Incidental, ambiguous or off-topic responses enter one denominator differently |
| Observation depth | The amount of usable evidence is visible and sufficient for the claim | A sparse topic is displayed with the same certainty as a well-observed topic |
| Time window | Collection periods are aligned | Topics are compared across materially different periods |
| Market and language | Relevant market and language conditions are aligned | Geography or language differs silently between the topic populations |
| Collection status | Both sets are complete enough for the comparison being made | Partial, unsupported or unavailable data are treated as complete measurements |
| Numerator and denominator | The underlying counts are visible | Percentages are compared without knowing how many observations produced them |
The contract is not a claim that every dimension must be mathematically identical.
It is a discipline for identifying differences that could materially change the interpretation.
If two topics fail a material part of the contract, the correct answer may be:
These results are not directly comparable yet.
That is more useful than manufacturing precision from two percentages.
03 · Prompt composition
How can prompt mix create an apparent topic gap?
Prompt composition can change the population being measured, so a raw difference between topic averages may partly reflect the kinds of questions assigned to each topic.
Kojable's AI Citations by Query Type research illustrates this problem.
The study grouped prompts by buyer-question type and examined visible citation exposure across ChatGPT, Gemini and Perplexity. Gemini showed the clearest raw category variation. But the largest category difference was strongly entangled with retrieval need. When the comparison was restricted to a common retrieval-need cohort, the observed difference became substantially smaller.
The practical lesson is not that retrieval need explains every visibility difference.
It is:
Before interpreting a topic difference, check whether the topics contain materially different kinds of questions.
Imagine two monitored topics.
Topic A is mostly made up of broad educational questions.
Topic B is mostly made up of current vendor comparisons, pricing questions and implementation requirements.
If Topic B produces more cited, grounded or evidence-heavy answers, it would be risky to attribute the difference to “Topic B performs better” without considering how the question populations differ.
The topic label is descriptive. It does not isolate every other factor associated with the observations inside that topic.
This is also why the same buyer question should not be represented by many trivial wording variations merely to increase sample size. A large collection of near-duplicate prompts may create more response rows without increasing the breadth of buyer decisions represented.
Kojable's research explicitly distinguishes response volume from prompt-template breadth. Repeating a small set of prompt designs does not provide the same generalisability as observing a broader set of distinct questions.
04 · Platform context
Why should platforms stay separate when comparing topics?
Platform differences can change the meaning of the same observed rate, so topic comparisons should normally preserve the platform dimension rather than silently averaging it away.
Kojable's retrieval-need research analysed visible citation behaviour across ChatGPT, Gemini and Perplexity. The same retrieval-need classification produced materially different observable patterns across the three environments.
These findings concern citation exposure specifically. They support keeping platform context visible during measurement; they do not establish that every AI visibility metric will vary across platforms in the same way.
The study's conclusion is not that one platform is inherently better or worse.
It is that the relationship being measured was platform-specific.
That matters for topic visibility.
Suppose Topic A contains proportionally more observations from a platform where the measured event is common, while Topic B contains more observations from a platform where the same event is less common or more variable.
A pooled Topic A versus Topic B percentage may then reflect:
- the topic;
- the platform mix;
- or both.
Without separating those dimensions, the team cannot know how much interpretation the topic comparison deserves.
A safer structure is:
Topic × buyer-question type × platform or surface
before moving to a portfolio-level average.
Kojable's Topic Visibility Monitoring similarly keeps portfolio, topic, question and platform or model evidence connected without treating those layers as interchangeable.
05 · Observation depth
Why doesn't a large response count automatically make a topic comparison stronger?
A large response count can increase observation depth without increasing the breadth of what has actually been tested.
Consider two hypothetical topic panels.
One contains 800 responses generated from a small number of closely related prompt templates.
The other contains fewer total responses but covers a wider range of genuinely distinct buyer questions.
The first panel has the larger response count. It does not automatically provide the broader picture of the topic.
This distinction appears directly in Kojable Research. The query-type study used a large response collection but still warns that repeated observations from the same prompt templates create clustering and that several narrower routes have insufficient breadth for broad generalisation.
The practical reporting rule is:
Show observation depth and prompt breadth separately.
Useful supporting fields include:
- eligible responses;
- distinct buyer questions;
- distinct prompt templates;
- platforms or surfaces represented;
- collection dates;
- and the relevant exclusion rules.
Do not replace all of those dimensions with one sample-size number.
There is also no universal evidence-backed number of prompts that makes every topic comparison valid.
The number required depends on what the topic represents, how heterogeneous the buyer questions are, how much run-to-run variation exists and how strong a conclusion the team intends to make.
A narrow operational topic and a broad category topic may require very different designs.
06 · Eligibility and denominators
How should relevance rules and denominators affect topic scores?
The denominator determines what the percentage actually means.
A topic can only be compared responsibly when its eligibility rules are understood.
Not every answer containing a target word is genuinely about the intended topic. A phrase may occur incidentally, belong to another industry, refer to another company or appear without enough context to classify confidently.
Kojable's current Topic Visibility Monitoring distinguishes between relevant, ambiguous and excluded observations before calculating customer-facing topic metrics. It also keeps the underlying observations available so teams can understand why something was included or excluded.
This matters because two topics can have identical-looking rates built from different eligibility decisions.
For an illustrative example:
- Topic A may have 100 observed answers, of which only 60 are genuinely relevant.
- Topic B may have 100 observed answers, with 95 genuinely relevant.
If the displayed metric silently uses all observed answers for one topic and relevant answers for the other, the percentages no longer answer the same question.
A defensible comparison requires both topics to use the same approved metric definition and eligibility rule. For a company-mention metric, for example:
eligible relevant responses containing the company ÷ all eligible relevant responses for the topic
The numerator and denominator should be visible enough for another analyst to understand what was counted.
Unsupported data create a related problem.
A topic with no valid observations is not the same as a topic measured at zero visibility.
Kojable's Topic Visibility Monitoring explicitly distinguishes sparse, unsupported, partially collected and unavailable states and warns that unavailable data should not be displayed as 0% visibility. Percentages should remain connected to their numerators, denominators, sample size and collection status.
That principle should apply to every topic-level comparison.
No evidence is not zero performance.
07 · Reporting strength
How should sample size and uncertainty change the language you use?
The strength of the wording should match the strength of the measurement.
A topic supported by broad, relevant and complete observations can support a stronger descriptive comparison than one based on sparse, partial or ambiguous evidence.
That does not require turning uncertainty into another opaque confidence score.
As a general comparison-design reference rather than evidence about AI visibility, NIST notes that sound comparisons depend on experimental design, repetition choices and uncertainty assessment.
Use explicit reporting states instead.
| Reporting state | What it means | Appropriate wording |
|---|---|---|
| Comparable | The major measurement dimensions are sufficiently aligned and evidence is adequate for the stated comparison | “Topic A showed lower visibility than Topic B in the tested panel.” |
| Comparable with qualification | A known difference remains but can be clearly disclosed or stratified | “Topic A was lower within the tested platform, although its prompt mix was narrower.” |
| Directional only | The evidence suggests a difference, but breadth or observation depth is too limited for a strong conclusion | “The current observations suggest weaker visibility for Topic A.” |
| Not comparable | A material design difference prevents a defensible direct comparison | “The current topic scores should not be directly compared.” |
| Insufficient evidence | There is not enough valid information to calculate or interpret the comparison | “There is insufficient evidence to assess Topic A.” |
These states are editorial reporting categories, not empirically validated statistical thresholds.
They are designed to prevent a sparse topic and a well-supported topic from receiving identical language simply because both have a percentage attached.
The same principle applies over time.
For illustration, a movement from 40% to 60% means little if the prompt set, platform mix, eligibility rules or measurement definition changed at the same time.
08 · Gap threshold
When should you call a difference a real topic-level visibility gap?
Call a topic difference a visibility gap only after the comparison is sufficiently defined and the observed difference survives the material comparability checks.
At minimum, the team should be able to answer:
- What exact visibility outcome are we comparing?
- Which buyer questions make up each topic?
- Are the platform or surface conditions aligned?
- Do the topics use the same relevance and eligibility rules?
- How many eligible observations and distinct prompt designs support each result?
- Were they collected over comparable periods?
- Is either topic sparse, partial or unsupported?
- Could a material difference in prompt composition explain part of the apparent gap?
If those questions are adequately answered, a conclusion such as this becomes defensible:
Topic A showed lower company mention visibility than Topic B across the comparable buyer-question and platform panel tested during the stated period.
Notice what the statement does not say.
It does not say:
- Topic A is universally weaker.
- Topic A needs more content.
- Competitor sources caused the difference.
- The AI system prefers Topic B.
- The gap is commercially important.
- A particular intervention should be prioritised.
Those are later questions.
This article's measurement job ends once the team has established that the difference itself is real enough to investigate.
09 · Diagnosis handoff
What happens after a comparable topic gap is established?
A comparable gap becomes an input to diagnosis. It is not a diagnosis by itself.
Once the team can defensibly say:
This buyer-relevant topic is repeatedly weaker under comparable measurement conditions
the next questions concern meaning and action:
- Does the pattern recur beyond the measured panel?
- Is the issue visibility, inaccurate representation, missing proof or another representation gap?
- Which observable evidence is associated with it?
- Does it affect an important buyer decision?
- Is there a realistic action path?
Those questions belong to Answer Intelligence, Kojable's diagnostic capability within the Diagnose stage.
Answer Intelligence deliberately separates direct observation from interpretation and asks whether a recurring pattern has enough evidence, commercial relevance and actionability to deserve intervention.
Keeping the stages separate prevents a measurement result from choosing its own explanation.
The workflow becomes:
Monitor → establish a comparable topic difference → Diagnose → Improve → Verify
A topic score tells you what was observed under defined conditions.
Diagnosis tells you what that observation may mean.
Quick answers
Frequently asked questions about topic-level AI visibility
What is topic-level AI visibility?
Topic-level AI visibility measures a company's presence or prominence within a defined topic or buyer-question cluster under a stated measurement design. The result should specify what was measured, which questions and platforms were included, which observations were eligible and what period the data covers.
Can you compare AI visibility scores across different topics?
Yes, but only when the scores measure sufficiently comparable populations or material differences are explicitly accounted for. Check the metric definition, buyer-question mix, platform or surface, relevance rules, observation depth, prompt breadth, time window and collection status before interpreting one topic as stronger or weaker.
Can a topic look weak because its prompts are different?
Yes. Different topics can contain different kinds of buyer questions, and those questions may have different information or retrieval needs. Kojable Research found that a large raw query-type citation difference on Gemini became much smaller once retrieval need was compared on more common support. That does not prove every topic gap is a prompt-composition effect, but it shows why prompt mix should be checked before interpreting the raw difference.
Should AI visibility from different platforms be averaged together?
Not automatically. Platform baselines and observable behaviour can differ. Kojable Research found materially different relationships between retrieval need and visible citation exposure across ChatGPT, Gemini and Perplexity. Preserve platform-level results first, then aggregate only when the pooled metric still answers a meaningful question.
How many prompts are needed for topic-level AI visibility?
There is no universal evidence-backed minimum. The required breadth depends on the topic, the diversity of buyer questions, the variability of the observed answers and the strength of the conclusion being made. Report distinct buyer questions or prompt templates separately from total response count because repeated observations do not automatically increase topic breadth.
Is no data the same as zero AI visibility?
No. A topic with no supported observations, partial collection or insufficient relevant responses should not automatically be displayed as zero visibility. The measurement state and underlying denominator should be reported explicitly.
What should you report when two topics cannot be compared?
State why the comparison is not defensible. For example, the topics may use different platform mixes, prompt types, eligibility rules or collection periods, or one topic may have insufficient evidence. Use a status such as not comparable, directional only or insufficient evidence rather than forcing a precise gap claim.
When does a topic-level difference become a gap worth diagnosing?
When the difference concerns a buyer-relevant visibility measure, is supported by sufficiently comparable measurement conditions and has enough valid observations to justify the descriptive claim. At that point the result can enter diagnosis. The measurement itself still does not establish why the gap exists or what should be changed.