Kojable research · Cross-Model Citation Study

Different Answers, Different Evidence: How Claude, Gemini, OpenAI and Perplexity Cite the Web

Published By Piush Vaish

The same buyer question can produce four plausible AI answers—and four largely different evidence sets.

Key finding

The four provider stacks frequently assembled materially different evidence environments around the same buyer questions. Average pairwise URL Jaccard overlap was only about 0.009–0.020, while visible citation density ranged from 5.28 citation events per 1,000 words for OpenAI to 39.79 for Gemini.

Qualification: This is a fixed observational benchmark with one observed run per provider-question cell. It does not identify a universal model preference, provider-wide accuracy ranking, causal citation mechanism or run-to-run stability.

  • Cross-provider AI citations
  • 10 designed questions
  • 39 successful responses
  • 14 min read
Editorial comparison of URL overlap among Claude, Gemini, OpenAI and Perplexity, showing average provider-pair Jaccard values between approximately 0.9% and 2.0% in the matched panel.
The same buyer questions often led the four provider stacks to cite largely different URL sets. The hero uses the locked URL-overlap values from the matched panel.
Open full-resolution figure
40expected provider-question cells
39successful observed responses
9questions in the primary matched panel
1 runper provider-question cellInterpretationRepeated runs are required to assess output stability.

Study overview

Executive summary

AI answers can look similar on the surface while being built from very different evidence underneath. Kojable presented the same ten B2B buyer questions to Claude, Gemini, OpenAI and Perplexity and analysed the citations, source mix, freshness, support relationships, publishers and authors behind the responses.

Of 40 expected provider-question cells, 39 succeeded. Gemini P08 returned no usable answer, so the primary four-provider comparison uses the nine questions completed by all four provider stacks.

  • Citation density

    Ranged from 5.28 citation events per 1,000 words for OpenAI to 39.79 for Gemini.

  • Cited-domain breadth

    Averaged 5.33 unique domains per question for Claude and 9.22 for Gemini at the observed range endpoints.

  • Source mix

    Observed independent-source share ranged from 22.0% for Perplexity to 65.1% for Claude.

  • Freshness

    Median cited-source age ranged from approximately 76 days for Claude to 120 days for Perplexity, conditional on publication-date coverage.

  • Citation placement

    Automated material-claim citation coverage ranged from 10.2% for OpenAI to 43.2% for Gemini.

  • Reviewed support

    Direct-support rates among reviewed relationships ranged from 72.0% for Gemini to 86.7% for OpenAI, but provider review coverage differed substantially.

  • Evidence overlap

    Pairwise URL overlap was extremely low: average within-question Jaccard values ranged from roughly 0.009 to 0.020.

  • Named authors

    Named-author share ranged from 50.9% for Gemini to 77.1% for Claude.

Answer first

The four provider stacks did not rely on the same evidence

Direct answer

Claude, Gemini, OpenAI and Perplexity differed substantially in visible citation density, cited-domain breadth, source mix and the individual URLs and publishers that appeared. The clearest signal was low source overlap: the same question often led the provider stacks to assemble largely different cited-source sets.

Representation in one provider's evidence environment cannot be assumed to transfer to another. Monitoring only one AI system can therefore miss part of the public information landscape associated with AI-mediated buyer research.

The study does not show that one provider is universally better, more accurate or more authoritative. It also does not establish that the observed source differences are stable across repeated runs. For definitions separating citations, retrieval, grounding and source influence, read Kojable's AI citation guide.

Finding 1

Citation intensity varied dramatically

Gemini produced the highest visible citation density in the matched panel at 39.79 citation events per 1,000 words. Perplexity followed at 21.06, Claude at 8.40, and OpenAI at 5.28.

Citation events per 1,000 response words for Claude, Gemini, OpenAI and Perplexity, with Gemini highest at 39.79 and OpenAI lowest at 5.28.
Figure 1. Citation events per 1,000 response words in the nine-question matched panel. Whiskers are benchmark-panel resampling intervals: they describe sensitivity to the composition of this fixed question panel, not run-to-run provider stability or population confidence intervals.
Open full-resolution figure

A citation event records that a source was attached to an answer placement. It does not establish that the source is authoritative, supports the claim correctly or makes the answer factually better. Nor does a larger number of visible citations necessarily mean that a provider consulted more evidence internally.

The defensible conclusion is narrower: under the same benchmark protocol, the provider stacks produced very different levels of visible citation activity. Raw citation counts should not be treated as directly interchangeable across systems.

Finding 2

Evidence breadth and concentration also differed

Gemini showed the broadest cited-domain profile, averaging 9.22 unique cited domains per question. OpenAI averaged 6.22, Perplexity 5.89, and Claude 5.33. Gemini and OpenAI had lower observed domain HHI values, while Claude and Perplexity were more concentrated.

Comparison of average unique cited domains per question and citation-event domain concentration across Claude, Gemini, OpenAI and Perplexity.
Figure 2. Average cited-domain breadth and citation-event concentration by provider in the matched panel. Wider breadth and lower concentration are descriptive characteristics; neither is treated as inherently better evidence.
Open full-resolution figure

A broader set can expose a reader to more perspectives, but breadth is not a quality score. A concentrated answer may rely on a small number of strong sources; a broad answer may contain weak, repetitive or commercially similar sources. Breadth and concentration therefore remain separate measurements.

Finding 3

The source ecosystems were materially different

Cited pages were classified across separate dimensions including ownership, commercial interest and content format. An independent source is a taxonomy category, not a guarantee of objectivity or factual quality. Commercially interested, competitor and first-party sources can all contain accurate and useful information.

Observed independent-source share ranged from 65.1% for Claude to 22.0% for Perplexity. Gemini was 34.2% and OpenAI 44.4%. The commercially interested-source dimension ranged from 34.9% for Claude to 83.6% for Perplexity.

Independent, commercially interested, competitor and first-party cited-source shares for Claude, Gemini, OpenAI and Perplexity.
Figure 3A. Observed cited-source relationship dimensions by provider. The dimensions are grouped rather than stacked because ownership and commercial-interest labels are separate taxonomy measures.
Open full-resolution figure
Heatmap comparing cited-page format shares across Claude, Gemini, OpenAI and Perplexity using the study's locked source-format taxonomy.
Figure 3B. Source-format composition by provider. The heatmap retains the locked format taxonomy; zero-only case-study rows are omitted from the display and no categories are combined.
Open full-resolution figure

The provider stacks drew on different mixtures of publishers, vendors, agencies, competitors, communities and first-party sources. AI representation is therefore not determined only by what exists on a company's own website.

  • First-party product information
  • Competitor narratives
  • Agency or consultancy content
  • Independent editorial coverage
  • Review, directory and community pages
  • Research or institutional material

Kojable's Monitor → Diagnose → Improve → Verify workflow separates the observed answer from the source and information gaps associated with it.

Finding 4

Cited-source freshness also differed

Source age was measured relative to the timestamp when each response was generated. Missing or ambiguous publication dates remain unknown rather than being assigned to an old-content bucket.

Median cited-source age and publication-date coverage
Provider Median cited-source age Publication-date coverage
Claude76 days88.8%
Gemini81 days64.8%
OpenAI88 days70.3%
Perplexity120 days97.2%
Median cited-source age and publication-date coverage for Claude, Gemini, OpenAI and Perplexity.
Figure 4. Cited-source age relative to each response-generation timestamp. Freshness estimates are conditional on valid publication dates; missing dates remain unknown rather than being treated as old.
Open full-resolution figure

These results do not show that newer content causes citations, that older pages are penalised or that companies should publish on a particular cadence. Page characteristics such as word count, headings, images, lists, tables and structured-data markers are also descriptive; turning them into causal rules requires an intervention study.

Finding 5

Citation quantity, placement and support are different questions

Citation intensity
How many visible citation events appear.
Material-claim citation coverage
How often automatically identified material claims have an observed citation placement.
Reviewed support
Whether a cited source directly or partially supports the linked claim in the reviewed sample.

Automated material-claim citation coverage was 43.2% for Gemini, 33.7% for Perplexity, 25.0% for Claude and 10.2% for OpenAI. For support quality, the study reviewed a stratified sample of 120 claim-citation relationships from 1,298 mapped relationships. Each sampled relationship was independently assessed by one human reviewer and one AI-assisted web-research reviewer, with disagreements adjudicated.

Direct support among rated links and annotation coverage
Provider Direct-support rate Annotation coverage
Claude78.5%19.5%
Gemini72.0%4.6%
OpenAI86.7%31.3%
Perplexity79.3%7.3%
Material-claim citation placement, direct and partial support, and annotation coverage across Claude, Gemini, OpenAI and Perplexity.
Figure 5. Automated material-claim citation placement and reviewed claim-to-source support. Direct-support percentages apply only to rated links and must be read with provider-specific annotation coverage; they are not provider-wide factual-accuracy rates.
Open full-resolution figure

Reviewed support metrics are not estimates of overall answer accuracy, hallucination rate or provider-wide factual reliability. Raw agreement between the two assessors was 77.5%, Cohen's kappa was approximately 0.477, and 27 disagreements required adjudication. A high number of citation markers does not automatically mean stronger evidentiary support.

Finding 6

Cross-provider evidence overlap was very low

Across provider pairs, average within-question URL Jaccard overlap ranged from approximately 0.009 to 0.020. Domain overlap was also low, with average Jaccard values roughly between 0.025 and 0.048. Publisher overlap remained low as well.

Provider-pair Jaccard overlap for cited URLs, domains and publishers across Claude, Gemini, OpenAI and Perplexity, with very low overlap throughout.
Figure 6. Average within-question source overlap across provider pairs. URL, domain and publisher overlap are reported separately. Low overlap shows that the observed provider stacks often assembled different cited-source sets; it does not establish that those differences are stable across repeated runs.
Open full-resolution figure

Answer similarity and evidence similarity are not the same. Two answers may reach similar conclusions from different publishers, domains or commercial ecosystems; two systems may also cite some of the same domains while emphasising different pages and claims.

Visibility in one provider's evidence environment does not guarantee visibility in another. Within this fixed benchmark, cross-provider evidence convergence was low, but repeated-run work is required to test stability.

Finding 7

Publisher and author visibility added another layer

A domain is not the same thing as a publisher, and a publisher is not the same thing as an author. The study resolved these entities separately where the available evidence allowed it.

Claude
77.1% named-author share
OpenAI
69.3% named-author share
Perplexity
65.5% named-author share
Gemini
50.9% named-author share
Named and verified author shares alongside author and publisher resolution coverage for Claude, Gemini, OpenAI and Perplexity.
Figure 7. Publisher and author visibility in cited-source records. Author and publisher resolution are coverage-qualified; named authorship is an observed characteristic, not evidence that bylines cause citation or establish authority.
Open full-resolution figure

The results do not prove that named authors are preferred or that bylines cause AI citation. They show authorship as another observable layer of the evidence environment. Unresolved identities remain unresolved rather than being guessed, and author-review provenance remains explicit.

Finding 8

Retrieval and citation are not the same stage

Claude and OpenAI exports expose candidate/search pools that the study treats as complete under its locked observability contract. This supports a distinction between sources observed in an exposed candidate pool and sources ultimately cited.

The same analysis is not valid for Gemini and Perplexity because their exposed source metadata is citation-biased rather than a validated complete search pool. Candidate-selection analysis is valid only where a sufficiently complete candidate/search pool was observable.

In the nine-question complete-case panel, the observed candidate-to-citation selection rate was approximately 11.0% for Claude and 6.2% for OpenAI.

Candidate-to-citation selection rates of approximately 11.0% for Claude and 6.2% for OpenAI, the two eligible provider surfaces.
Figure 8. Candidate-to-citation selection for Claude and OpenAI, the two provider stacks with candidate pools treated as complete under the study contract. The selection fraction is an operational diagnostic, not a provider-quality score, and is not extrapolated to Gemini or Perplexity.
Open full-resolution figure

The difference does not mean Claude is better at selection. A broader candidate pool can naturally produce a lower selection fraction, and the denominator is partly a property of retrieval and instrumentation.

  1. Candidate observation

    Was the source observed in an eligible exposed candidate pool?

  2. Citation selection

    If observed, was the source selected for citation?

  3. Claim relationship

    If cited, which claim or section did it support, and did the cited page support that claim?

  4. Provider comparison

    How did those observable states differ across eligible provider surfaces?

Practical interpretation

What these findings mean for companies

The study supports treating AI representation as a multi-provider evidence problem, not only as an answer-ranking or mention-count problem. A company can be represented differently because provider stacks can assemble information from different parts of the web.

  • Observed visibility

    Which pages about the company appear, and which are selected into final answers across provider surfaces?

  • Narrative environment

    Which third-party publishers, authors and competitor pages shape the category narrative?

  • Evidence gaps

    Are important proof points absent, outdated claims still visible or accurate first-party facts difficult to corroborate?

  • Provider-specific patterns

    Does a provider rely on a narrow domain set or draw from a materially different evidence mix?

The study does not provide a universal recipe for getting cited by AI. It supports a measurement principle: look beyond the mention and inspect the evidence environment around the answer. Kojable's AI citation monitoring applies that distinction to company-specific buyer questions.

Method

Study design

Designed panel
Ten B2B buyer questions across Claude, Gemini, OpenAI and Perplexity: 40 expected provider-question cells.
Successful responses
Thirty-nine cells succeeded. Gemini P08 returned no usable answer.
Primary matched panel
Nine questions completed by all providers: P01, P02, P03, P04, P05, P06, P07, P09 and P10.
Analytical unit
The question is primary. Citation events, URLs and source records are nested within provider-question responses. Each provider-question cell contains one observed run.
Provider summaries
Equal-weight question-level macro averages are used for primary provider summaries.
Resampling
Matched-question Poisson block resampling applies the same question-level weight across eligible providers within a replicate.

Provider observability

Provider-native source metadata does not expose the same object on every surface. The eligibility distinction is part of the analysis, not a missing-data convenience.

Candidate-pool observability and selection-analysis eligibility
Provider Exposed candidate pool Candidate-selection analysis
ClaudeTreated as completeEligible
OpenAITreated as completeEligible
GeminiCitation-biasedNot eligible
PerplexityCitation-biasedNot eligible

A source absent from Gemini or Perplexity's exposed metadata cannot be described as “not retrieved”; their underlying candidate pools are not fully observable in the current exports.

Evidence boundaries

Limitations

  • Designed question panel

    The ten questions were designed for this study rather than randomly sampled from all possible B2B buyer questions.

  • One observed run per cell

    The study does not estimate run-to-run volatility or longitudinal stability.

  • Instructed protocol

    The research instruction may influence search behaviour, source selection, answer structure and citation behaviour.

  • Provider-stack comparison

    Results can combine model behaviour, search backend, query generation, retrieval orchestration, API instrumentation, citation implementation and answer style. The study does not isolate those components causally.

  • Unequal observability

    Candidate-to-citation selection is restricted to Claude and OpenAI. Gemini and Perplexity do not expose validated complete candidate pools under the study contract.

  • Coverage-qualified enrichment

    Publication-date, page-feature, publisher and author resolution are incomplete to varying degrees. Missing values remain unknown rather than absent or zero.

  • Stratified support review

    Support review covers a subset of mapped claim-citation relationships and uses one human plus one AI-assisted reviewer. It is not a provider-wide accuracy estimate.

Most importantly, this is one snapshot of probabilistic systems. Repeated-run experiments are required before treating the observed differences as stable provider traits.

Research record

Research and reproducibility

The publication package makes the reported numbers traceable through the locked analysis contract, provider-observability rules, metric definitions, taxonomy and enrichment methods, provider scorecards, publication tables, benchmark-panel resampling outputs, claim-level traceability, and input/output manifests and checksums.

The publication layer does not create a composite provider score, declare an overall winner or introduce a causal claim about why the provider stacks differed.

View the Cross-Model Citation Study repository on GitHub.

FAQ

Frequently asked questions

Do Claude, Gemini, OpenAI and Perplexity cite the same sources?

Not usually in this fixed benchmark. Average pairwise URL Jaccard overlap across the matched questions was approximately 0.009–0.020, indicating that the observed provider stacks frequently cited largely different URL sets for the same buyer questions.

Which provider produced the most citations?

Gemini had the highest visible citation density in the matched panel at 39.79 citation events per 1,000 words. Perplexity was 21.06, Claude 8.40 and OpenAI 5.28. Citation density is a placement metric, not a measure of answer quality or factual accuracy.

Does producing more citations mean an AI answer is better supported?

No. Citation volume, material-claim citation placement and reviewed claim-to-source support are different measurements. A high citation count does not by itself establish that the sources directly support the claims or that the overall answer is more accurate.

Which provider cited the highest share of independent sources?

Claude had the highest observed independent-source share in the matched panel at 65.1%. Perplexity had the lowest at 22.0%. “Independent” is a source-taxonomy category and is not automatically a judgment of authority, objectivity or accuracy.

Why is candidate-to-citation selection reported only for Claude and OpenAI?

Under the study's observability contract, Claude and OpenAI expose candidate/search pools that are treated as complete for this analysis. Gemini and Perplexity expose citation-biased source metadata, so an equivalent retrieval-to-citation selection rate would not be valid.

Does this study show stable provider preferences?

No. Each provider-question cell contains one observed run. The benchmark shows differences in the recorded provider stacks under the tested protocol, but repeated runs are needed before those differences can be treated as stable provider traits.

From benchmark to company evidence

See what this looks like for your company

Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.

See how Kojable works →
Explore AI citation monitoring →

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish