Kojable research

Retrieval need is associated with citation exposure only on Gemini

Published By Piush Vaish

A three-platform analysis found a 45.2 percentage-point High-minus-Low difference on Gemini, negligible positive separation on ChatGPT and a ceiling effect on Perplexity.

Every Low-need Gemini response in the available data belonged to the Educational route. The analysis identifies an association, not a causal mechanism.

Retrieval need was associated with citation exposure on Gemini, but not as a general cross-platform rule. The evidence does not show that retrieval need caused citations or that the relationship follows a smooth Low-to-Medium-to-High gradient.

  • AI citation exposure
  • Retrieval need
  • Three platforms
  • 10 min read
Diagram showing one citation metric branching into several different platform behaviours.
55,315 responses analysed
3 included platforms
984 shared prompt templates
+45.2 points Gemini High minus Low

Qualification: Every Low-need Gemini response in the available data belonged to the Educational route. The analysis identifies an association, not a causal mechanism.

Study overview

Executive summary

A single citation-rate metric can hide materially different platform behaviours.

In an analysis of 55,315 responses from ChatGPT, Gemini and Perplexity, Gemini’s High-need prompts had a 91.3% citation-exposure rate, compared with 46.1% for Low-need prompts. The observed High-minus-Low difference was 45.2 percentage points, with a conservative 95% interval from 40.3 to 50.1 points.

ChatGPT showed no material positive relationship. Its High-minus-Low estimate was -2.2 points, with an interval from -3.9 to +0.3. Perplexity’s estimate was -0.1 points, with citation exposure close to 100% at every retrieval-need level.

The Gemini result is narrower than a causal headline would suggest. High and Medium citation rates were close, so the pattern was mainly a Low-versus-rest discontinuity. Every Low-need Gemini response in the available data also belonged to the Educational route. The analysis cannot separate retrieval need from route composition, repeated prompt-template effects or their interaction.

For marketing leaders, the implication is practical: segment AI representation by platform and commercially relevant buyer question before choosing a content, PR, SEO, brand or evidence action.

Answer first

Direct answer

Retrieval need was associated with citation exposure on Gemini, but not as a general cross-platform rule.

The evidence does not show that retrieval need caused citations or that the relationship follows a smooth Low-to-Medium-to-High gradient.

Branching diagram illustrating that retrieval need was associated with citation exposure only on Gemini.

Research question

Why this question matters

AI citation dashboards can create an impression of comparability. A percentage appears precise, but it may combine different models, prompt types and outcome conditions.

That is a problem for a CMO because the metric can influence where work is assigned. A perceived citation gap may be sent to content, SEO, PR or product marketing before the team knows whether the gap is recurring, model-specific, tied to a buyer question or connected to a weak source.

The first analytical question should therefore be narrower:

Does visible citation exposure change with a prompt-derived measure of retrieval need, and does that relationship differ by platform?

The study tests an observable proxy. It does not observe a platform’s internal retrieval trigger.

Study design

What was measured

The analysis contains 55,315 responses from ChatGPT, Gemini and Perplexity. One platform pending collection validation was excluded from Decision 1. Three responses had Unknown retrieval need and were excluded from inference and visualisations.

Each prompt was classified as High, Medium or Low retrieval need using its final prompt taxonomy. The outcome, has_raw_citations, records whether at least one citation appeared in the final response.

The known-need cohort contained:

  • 39,913 High-need responses, representing 72.2% of the cohort
  • 14,012 Medium-need responses, representing 25.3%
  • 1,387 Low-need responses, representing 2.5%

For each platform and need level, the analysis calculated an observed citation rate and a 95% Wilson interval. The High-minus-Low comparison used a conservative interval constructed from the Wilson bounds. Cramér’s V was used as a practical association measure.

The same 984 prompt templates were represented on all three included platforms. This reduces the possibility that the main cross-platform pattern came from completely unmatched template coverage.

Figure 1. Citation exposure by retrieval need and platform. ChatGPT remains high with negligible separation, Gemini shows a Low-versus-rest discontinuity, and Perplexity remains near citation saturation.
Open full-resolution figure

Platform results

Three platforms showed three regimes

Citation exposure and High-minus-Low differences across the three included product environments.
Platform High-need citation exposure Medium-need citation exposure Low-need citation exposure High-minus-Low difference Interpretation
ChatGPT 94.4% (n=12,019) 94.5% (n=4,235) 96.6% (n=440) -2.2 percentage points (95% interval: -3.9 to +0.3) No material positive relationship; Cramér’s V was 0.015.
Gemini 91.3% (n=14,253) 89.9% (n=4,971) 46.1% (n=484) +45.2 percentage points (95% interval: +40.3 to +50.1) A Low-versus-rest separation, not a smooth gradient; Cramér’s V was 0.231.
Perplexity 99.9% (n=13,641) 99.9% (n=4,806) 100.0% (n=463) -0.1 percentage points (95% interval: -0.2 to +0.7) A ceiling effect left almost no variation for the binary outcome to explain.

Gemini qualification: Every Low-need Gemini response belonged to the Educational route. The +45.2-point difference is an association and does not identify a causal mechanism.

ChatGPT: high citation exposure with negligible separation

ChatGPT’s citation-exposure rates were:

  • High need: 94.4%, n=12,019
  • Medium need: 94.5%, n=4,235
  • Low need: 96.6%, n=440

The High-minus-Low estimate was -2.2 percentage points. Its conservative interval ran from -3.9 to +0.3 points.

Cramér’s V was 0.015, and the reported chi-square test was not statistically significant at conventional levels. The data provide no evidence that greater retrieval need increased ChatGPT citation exposure in this cohort.

Gemini: a large Low-versus-rest separation

Gemini’s citation-exposure rates were:

  • High need: 91.3%, n=14,253
  • Medium need: 89.9%, n=4,971
  • Low need: 46.1%, n=484

The High-minus-Low difference was +45.2 percentage points, with a conservative interval from +40.3 to +50.1 points. Cramér’s V was 0.231, the only meaningful practical association reported among the three platforms.

Nearby qualification: Every Low-need Gemini response belonged to the Educational route. The result is an association and does not identify a causal mechanism.

The more precise interpretation is not “citations increased steadily with retrieval need”. High was only 1.4 points above Medium. The dominant pattern was the unusually low citation exposure of the Low-need branch.

Perplexity: a ceiling effect

Perplexity’s rates were:

  • High need: 99.9%, n=13,641
  • Medium need: 99.9%, n=4,806
  • Low need: 100.0%, n=463

The High-minus-Low difference was -0.1 points, with a conservative interval from -0.2 to +0.7.

This is not evidence that Perplexity necessarily used the same internal retrieval process for every prompt. It means that a binary visible-citation outcome had almost no remaining variation for retrieval need to explain.

Figure 2. High-minus-Low citation-rate differences. Only Gemini shows a material positive separation in the observed cohort.

Every Low-need Gemini response in the available data belonged to the Educational route; the +45.2-point estimate does not establish causation.

Open full-resolution figure

Important qualification

The 45.2-point result needs a nearby qualification

The Gemini result is large, but the study design does not identify a causal mechanism.

Every Low-need Gemini response in the available data belonged to the Educational route. There is no Low-need overlap with other Gemini routes in this branch.

The observed separation may therefore reflect:

  • the retrieval-need classification
  • Educational-route composition
  • a cluster of repeated prompt templates
  • an interaction among those factors

The study detects the association. It cannot determine which factor caused it.

This limitation should remain visible wherever the 45.2-point number appears. Moving it to a footnote would make the headline stronger than the evidence.

Template analysis

Shared templates reduce one concern and reveal another

The same 984 templates appeared on all three platforms. The median template citation rate was 100% for ChatGPT, Gemini and Perplexity.

At a 90% citation-rate benchmark:

  • 92.0% of ChatGPT templates met or exceeded the threshold
  • 87.5% of Gemini templates met or exceeded it
  • 99.9% of Perplexity templates met or exceeded it

Most Gemini templates therefore still had very high citation exposure. The platform-level difference was concentrated in a lower tail rather than spread uniformly across all prompt designs.

Equal-template weighting reduces the influence of frequently repeated templates. It does not make templates exchangeable, remove within-template dependence or prove causal retrieval behaviour.

Figure 3. Template-level citation-rate survival curves. At a 90% citation-rate benchmark, 92.0% of ChatGPT templates, 87.5% of Gemini templates and 99.9% of Perplexity templates met or exceeded the threshold.
Open full-resolution figure

Diagnostic boundary

What citation exposure does not measure

Citation exposure means that at least one visible citation appeared in the final answer.

It does not establish:

  • whether retrieval occurred without a visible citation
  • whether the cited source supports the answer
  • whether the source is authoritative or current
  • whether the answer represents a company accurately
  • whether a source caused the answer
  • whether citation presence affects buyer consideration or commercial outcomes

Those are separate diagnostic questions. Kojable’s citation and source analysis treats citation presence as a signal to investigate, not a diagnosis by itself.

Layered diagnostic diagram showing citation presence as one signal among source quality, answer support and positioning.

For marketing leaders

What a CMO should do differently

  1. Separate the baseline by model

    A pooled citation rate can conceal different platform regimes. Review ChatGPT, Gemini, Perplexity and other in-scope systems separately before combining results. AI representation monitoring should preserve that platform context.

  2. Organise prompts around buyer questions

    Category discovery, comparison, trust, pricing and general education may behave differently. Grouping them only by total volume can hide the context that determines the next action. The AI answer alignment workflow starts from relevant buyer questions.

  3. Separate citation presence from citation usefulness

    For each recurring source, assess:

    • relevance to the claim
    • authority
    • recency
    • ownership
    • actionability
    • whether it supports the wording in the answer
  4. Diagnose before prescribing content

    A lower citation rate does not automatically imply a content gap. The issue may be stale pricing, unclear positioning, missing proof, weak documentation, competitor framing or a third-party evidence gap. A free Brand Integrity Audit can establish the baseline before work is assigned.

  5. Retest comparable questions

    After a priority improvement is carried out, retest comparable prompts and assess what changed. Results may be clear, mixed, delayed or inconclusive. No company controls a third-party AI system.

Operating process

From metric to operating process

Kojable is an AI answer alignment platform for B2B companies. It connects four stages:

Monitor

Establish the current representation across relevant models, buyer questions, competitors and source contexts.

Diagnose

Identify recurring claims, citation patterns, missing proof, outdated information and realistic action paths.

Improve

Prioritise what should change, why, where and how.

Verify

Retest comparable questions and assess what moved, what held and what needs further attention.

The aim is not to replace every visibility, SEO, content or PR tool. It is to connect observed signals to an evidence-backed decision and a comparable retest.

Study boundaries

Limitations

This is an observational analysis of a prompt-derived proxy and a visible binary outcome.

The retrieval-need label is not a record of what a platform internally judged. Repeated responses from the same template create dependence that row-level Wilson intervals do not model directly. The Low-need group represents 2.5% of the known-need cohort and is compositionally narrow. Perplexity’s near-universal citation exposure creates a ceiling effect.

Website source note: Kojable analysis of 55,315 responses from ChatGPT, Gemini and Perplexity. Retrieval need was classified from prompt content. Citation exposure records whether at least one visible citation appeared. Three Unknown-need rows were excluded from inference and visualisations.

Method and governance

Research and reproducibility details

Collection period and scope

Data were collected between 10 September 2025 and 17 June 2026, inclusive.

The platform-specific collection windows for Decision 1 were:

  • ChatGPT: 10 September 2025–17 June 2026
  • Gemini: 12 September 2025–17 June 2026
  • Perplexity: 10 September 2025–17 June 2026

The analysed collection used English-language prompts for the United States market.

Product environments and model identification

Scrunch AI identified the product environments as ChatGPT, Gemini and Perplexity but did not expose the exact underlying model names or versions.

Results are therefore reported at the product-environment level. They must not be attributed to a particular underlying model release.

The analysis does not guess model names, infer model versions from collection dates or imply that the same underlying model was used throughout the collection window.

Retrieval-need taxonomy

Prompts were classified as High, Medium or Low retrieval need using an analysis-created, rule-based taxonomy derived from prompt content. Automated Unknown classifications were reviewed and manually labelled where a defensible High, Medium or Low assignment could be made.

Three responses remained Unknown in the final dataset and were excluded from inference and visualisations.

The retrieval-need classification is a prompt-derived analytical proxy. It does not record whether an AI platform internally decided to retrieve external information.

Citation-detection method

Citation exposure was measured using the Scrunch AI has_raw_citations field. A response was classified as having citation exposure when Scrunch AI recorded at least one raw citation associated with the final response.

This binary outcome does not establish:

  • whether retrieval occurred without a visible citation
  • whether the cited source supported the answer
  • whether the source was authoritative or current
  • whether the answer represented a company accurately
  • whether a source caused the answer
  • whether citation presence affected commercial outcomes

Excluded-platform handling

Claude was present in the wider collection programme but was excluded from Decision 1 because its citation data required additional validation and could not be treated as sufficiently reliable for this analysis.

No Claude responses were included in the reported sample counts, platform-level estimates, confidence intervals, association measures or figures.

Reproducibility and sign-off

No public reproducibility package accompanies this publication.

The final analysis, figures, interpretation and methodological limitations were reviewed and approved for publication by Piush Vaish, founder and CEO of Kojable, on 29 July 2026.

This was an author review and publication sign-off rather than an independent external review or independent replication.

FAQ

Frequently asked questions

Does retrieval need cause Gemini to cite more often?

No causal conclusion is supported. The study found an association between a prompt-derived retrieval-need classification and visible citation exposure on Gemini.

Is the Gemini result a smooth gradient?

No. High and Medium citation rates were close. The strongest separation was between Low need and the other two groups.

Why did Perplexity show no difference?

Perplexity cited almost every response at every need level. This ceiling left little variation for a binary citation-exposure outcome to explain.

Does a visible citation mean the answer is well supported?

Not necessarily. Citation presence does not measure source authority, relevance, claim support, recency or answer accuracy.

What should a marketing team measure next?

Segment by model and buyer question, examine the cited sources and recurring claims, identify the action owner, then retest comparable questions after priority work is completed.

Author

About the author

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. His experience spans high-growth technology companies and large enterprise environments. He combines technical depth with customer discovery, creative problem-solving and a strong bias towards shipping useful products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish

Principal CTA

Establish your own platform-specific baseline

Run the free Brand Integrity Audit to establish a platform- and buyer-question-specific baseline for how AI currently represents and cites your company.

Run the free Brand Integrity Audit