Kojable research
Retrieval need is associated with citation exposure only on Gemini
A three-platform analysis found a 45.2 percentage-point High-minus-Low difference on Gemini, negligible positive separation on ChatGPT and a ceiling effect on Perplexity.
Every Low-need Gemini response in the available data belonged to the Educational route. The analysis identifies an association, not a causal mechanism.
Retrieval need was associated with citation exposure on Gemini, but not as a general cross-platform rule. The evidence does not show that retrieval need caused citations or that the relationship follows a smooth Low-to-Medium-to-High gradient.
Qualification: Every Low-need Gemini response in the available data belonged to the Educational route. The analysis identifies an association, not a causal mechanism.
Study overview
Executive summary
A single citation-rate metric can hide materially different platform behaviours.
In an analysis of 55,315 responses from ChatGPT, Gemini and Perplexity, Gemini’s High-need prompts had a 91.3% citation-exposure rate, compared with 46.1% for Low-need prompts. The observed High-minus-Low difference was 45.2 percentage points, with a conservative 95% interval from 40.3 to 50.1 points.
ChatGPT showed no material positive relationship. Its High-minus-Low estimate was -2.2 points, with an interval from -3.9 to +0.3. Perplexity’s estimate was -0.1 points, with citation exposure close to 100% at every retrieval-need level.
The Gemini result is narrower than a causal headline would suggest. High and Medium citation rates were close, so the pattern was mainly a Low-versus-rest discontinuity. Every Low-need Gemini response in the available data also belonged to the Educational route. The analysis cannot separate retrieval need from route composition, repeated prompt-template effects or their interaction.
For marketing leaders, the implication is practical: segment AI representation by platform and commercially relevant buyer question before choosing a content, PR, SEO, brand or evidence action.
Answer first
Direct answer
Retrieval need was associated with citation exposure on Gemini, but not as a general cross-platform rule.
The evidence does not show that retrieval need caused citations or that the relationship follows a smooth Low-to-Medium-to-High gradient.
Research question
Why this question matters
AI citation dashboards can create an impression of comparability. A percentage appears precise, but it may combine different models, prompt types and outcome conditions.
That is a problem for a CMO because the metric can influence where work is assigned. A perceived citation gap may be sent to content, SEO, PR or product marketing before the team knows whether the gap is recurring, model-specific, tied to a buyer question or connected to a weak source.
The first analytical question should therefore be narrower:
Does visible citation exposure change with a prompt-derived measure of retrieval need, and does that relationship differ by platform?
The study tests an observable proxy. It does not observe a platform’s internal retrieval trigger.
Study design
What was measured
The analysis contains 55,315 responses from ChatGPT, Gemini and Perplexity. One platform pending collection validation was excluded from Decision 1. Three responses had Unknown retrieval need and were excluded from inference and visualisations.
Each prompt was classified as High, Medium or Low retrieval need using its final prompt taxonomy. The outcome, has_raw_citations, records whether at least one citation appeared in the final response.
The known-need cohort contained:
- 39,913 High-need responses, representing 72.2% of the cohort
- 14,012 Medium-need responses, representing 25.3%
- 1,387 Low-need responses, representing 2.5%
For each platform and need level, the analysis calculated an observed citation rate and a 95% Wilson interval. The High-minus-Low comparison used a conservative interval constructed from the Wilson bounds. Cramér’s V was used as a practical association measure.
The same 984 prompt templates were represented on all three included platforms. This reduces the possibility that the main cross-platform pattern came from completely unmatched template coverage.
Platform results
Three platforms showed three regimes
| Platform | High-need citation exposure | Medium-need citation exposure | Low-need citation exposure | High-minus-Low difference | Interpretation |
|---|---|---|---|---|---|
| ChatGPT | 94.4% (n=12,019) | 94.5% (n=4,235) | 96.6% (n=440) | -2.2 percentage points (95% interval: -3.9 to +0.3) | No material positive relationship; Cramér’s V was 0.015. |
| Gemini | 91.3% (n=14,253) | 89.9% (n=4,971) | 46.1% (n=484) | +45.2 percentage points (95% interval: +40.3 to +50.1) | A Low-versus-rest separation, not a smooth gradient; Cramér’s V was 0.231. |
| Perplexity | 99.9% (n=13,641) | 99.9% (n=4,806) | 100.0% (n=463) | -0.1 percentage points (95% interval: -0.2 to +0.7) | A ceiling effect left almost no variation for the binary outcome to explain. |
Gemini qualification: Every Low-need Gemini response belonged to the Educational route. The +45.2-point difference is an association and does not identify a causal mechanism.
ChatGPT: high citation exposure with negligible separation
ChatGPT’s citation-exposure rates were:
- High need: 94.4%, n=12,019
- Medium need: 94.5%, n=4,235
- Low need: 96.6%, n=440
The High-minus-Low estimate was -2.2 percentage points. Its conservative interval ran from -3.9 to +0.3 points.
Cramér’s V was 0.015, and the reported chi-square test was not statistically significant at conventional levels. The data provide no evidence that greater retrieval need increased ChatGPT citation exposure in this cohort.
Gemini: a large Low-versus-rest separation
Gemini’s citation-exposure rates were:
- High need: 91.3%, n=14,253
- Medium need: 89.9%, n=4,971
- Low need: 46.1%, n=484
The High-minus-Low difference was +45.2 percentage points, with a conservative interval from +40.3 to +50.1 points. Cramér’s V was 0.231, the only meaningful practical association reported among the three platforms.
Nearby qualification: Every Low-need Gemini response belonged to the Educational route. The result is an association and does not identify a causal mechanism.
The more precise interpretation is not “citations increased steadily with retrieval need”. High was only 1.4 points above Medium. The dominant pattern was the unusually low citation exposure of the Low-need branch.
Perplexity: a ceiling effect
Perplexity’s rates were:
- High need: 99.9%, n=13,641
- Medium need: 99.9%, n=4,806
- Low need: 100.0%, n=463
The High-minus-Low difference was -0.1 points, with a conservative interval from -0.2 to +0.7.
This is not evidence that Perplexity necessarily used the same internal retrieval process for every prompt. It means that a binary visible-citation outcome had almost no remaining variation for retrieval need to explain.
Every Low-need Gemini response in the available data belonged to the Educational route; the +45.2-point estimate does not establish causation.
Open full-resolution figureImportant qualification
The 45.2-point result needs a nearby qualification
The Gemini result is large, but the study design does not identify a causal mechanism.
Every Low-need Gemini response in the available data belonged to the Educational route. There is no Low-need overlap with other Gemini routes in this branch.
The observed separation may therefore reflect:
- the retrieval-need classification
- Educational-route composition
- a cluster of repeated prompt templates
- an interaction among those factors
The study detects the association. It cannot determine which factor caused it.
This limitation should remain visible wherever the 45.2-point number appears. Moving it to a footnote would make the headline stronger than the evidence.
Diagnostic boundary
What citation exposure does not measure
Citation exposure means that at least one visible citation appeared in the final answer.
It does not establish:
- whether retrieval occurred without a visible citation
- whether the cited source supports the answer
- whether the source is authoritative or current
- whether the answer represents a company accurately
- whether a source caused the answer
- whether citation presence affects buyer consideration or commercial outcomes
Those are separate diagnostic questions. Kojable’s citation and source analysis treats citation presence as a signal to investigate, not a diagnosis by itself.
For marketing leaders
What a CMO should do differently
-
Separate the baseline by model
A pooled citation rate can conceal different platform regimes. Review ChatGPT, Gemini, Perplexity and other in-scope systems separately before combining results. AI representation monitoring should preserve that platform context.
-
Organise prompts around buyer questions
Category discovery, comparison, trust, pricing and general education may behave differently. Grouping them only by total volume can hide the context that determines the next action. The AI answer alignment workflow starts from relevant buyer questions.
-
Separate citation presence from citation usefulness
For each recurring source, assess:
- relevance to the claim
- authority
- recency
- ownership
- actionability
- whether it supports the wording in the answer
-
Diagnose before prescribing content
A lower citation rate does not automatically imply a content gap. The issue may be stale pricing, unclear positioning, missing proof, weak documentation, competitor framing or a third-party evidence gap. A free Brand Integrity Audit can establish the baseline before work is assigned.
-
Retest comparable questions
After a priority improvement is carried out, retest comparable prompts and assess what changed. Results may be clear, mixed, delayed or inconclusive. No company controls a third-party AI system.
Operating process
From metric to operating process
Kojable is an AI answer alignment platform for B2B companies. It connects four stages:
Monitor
Establish the current representation across relevant models, buyer questions, competitors and source contexts.
Diagnose
Identify recurring claims, citation patterns, missing proof, outdated information and realistic action paths.
Improve
Prioritise what should change, why, where and how.
Verify
Retest comparable questions and assess what moved, what held and what needs further attention.
The aim is not to replace every visibility, SEO, content or PR tool. It is to connect observed signals to an evidence-backed decision and a comparable retest.
Study boundaries
Limitations
This is an observational analysis of a prompt-derived proxy and a visible binary outcome.
The retrieval-need label is not a record of what a platform internally judged. Repeated responses from the same template create dependence that row-level Wilson intervals do not model directly. The Low-need group represents 2.5% of the known-need cohort and is compositionally narrow. Perplexity’s near-universal citation exposure creates a ceiling effect.
Website source note: Kojable analysis of 55,315 responses from ChatGPT, Gemini and Perplexity. Retrieval need was classified from prompt content. Citation exposure records whether at least one visible citation appeared. Three Unknown-need rows were excluded from inference and visualisations.
Method and governance
Research and reproducibility details
Collection period and scope
Data were collected between 10 September 2025 and 17 June 2026, inclusive.
The platform-specific collection windows for Decision 1 were:
- ChatGPT: 10 September 2025–17 June 2026
- Gemini: 12 September 2025–17 June 2026
- Perplexity: 10 September 2025–17 June 2026
The analysed collection used English-language prompts for the United States market.
Product environments and model identification
Scrunch AI identified the product environments as ChatGPT, Gemini and Perplexity but did not expose the exact underlying model names or versions.
Results are therefore reported at the product-environment level. They must not be attributed to a particular underlying model release.
The analysis does not guess model names, infer model versions from collection dates or imply that the same underlying model was used throughout the collection window.
Retrieval-need taxonomy
Prompts were classified as High, Medium or Low retrieval need using an analysis-created, rule-based taxonomy derived from prompt content. Automated Unknown classifications were reviewed and manually labelled where a defensible High, Medium or Low assignment could be made.
Three responses remained Unknown in the final dataset and were excluded from inference and visualisations.
The retrieval-need classification is a prompt-derived analytical proxy. It does not record whether an AI platform internally decided to retrieve external information.
Citation-detection method
Citation exposure was measured using the Scrunch AI has_raw_citations field. A response was classified as having citation exposure when Scrunch AI recorded at least one raw citation associated with the final response.
This binary outcome does not establish:
- whether retrieval occurred without a visible citation
- whether the cited source supported the answer
- whether the source was authoritative or current
- whether the answer represented a company accurately
- whether a source caused the answer
- whether citation presence affected commercial outcomes
Excluded-platform handling
Claude was present in the wider collection programme but was excluded from Decision 1 because its citation data required additional validation and could not be treated as sufficiently reliable for this analysis.
No Claude responses were included in the reported sample counts, platform-level estimates, confidence intervals, association measures or figures.
Reproducibility and sign-off
No public reproducibility package accompanies this publication.
The final analysis, figures, interpretation and methodological limitations were reviewed and approved for publication by Piush Vaish, founder and CEO of Kojable, on 29 July 2026.
This was an author review and publication sign-off rather than an independent external review or independent replication.
FAQ
Frequently asked questions
Does retrieval need cause Gemini to cite more often?
No causal conclusion is supported. The study found an association between a prompt-derived retrieval-need classification and visible citation exposure on Gemini.
Is the Gemini result a smooth gradient?
No. High and Medium citation rates were close. The strongest separation was between Low need and the other two groups.
Why did Perplexity show no difference?
Perplexity cited almost every response at every need level. This ceiling left little variation for a binary citation-exposure outcome to explain.
Does a visible citation mean the answer is well supported?
Not necessarily. Citation presence does not measure source authority, relevance, claim support, recency or answer accuracy.
What should a marketing team measure next?
Segment by model and buyer question, examine the cited sources and recurring claims, identify the action owner, then retest comparable questions after priority work is completed.
Principal CTA
Establish your own platform-specific baseline
Run the free Brand Integrity Audit to establish a platform- and buyer-question-specific baseline for how AI currently represents and cites your company.
Run the free Brand Integrity Audit