Kojable research · Cross-Model Citation Study

Do Claude, Gemini, OpenAI and Perplexity Cite the Same Sources?

Published Updated By Piush Vaish

Four AI systems can answer the same buyer question without relying on the same evidence.

Key finding

Average within-question URL Jaccard overlap ranged from approximately 0.9% to 2.0% across the six provider pairs, while domain and publisher overlap remained below 5% across every provider pair.

Qualification: This companion analysis uses the same fixed panel and provider-question responses as the flagship Cross-Model Citation Study. It is not a separate experiment or an independent replication, and each provider-question cell contains one observed run.

  • 4 provider stacks
  • 9 matched questions
  • 6 provider pairs
  • 10 min read
Editorial comparison of URL overlap among Claude, Gemini, OpenAI and Perplexity, showing average provider-pair Jaccard values between approximately 0.9% and 2.0% in the matched panel.
The same buyer questions often led the four provider stacks to cite largely different URL sets. This reused flagship visual is based on the locked matched-panel URL-overlap values.
Open full-resolution figure
10designed B2B buyer questions
39 of 40successful provider-question responses
9questions in the primary matched panel
1 runper provider-question cellInterpretationRepeated runs are required to assess provider stability.

Answer first

Do the four systems cite the same sources?

Answer

No—not often in this benchmark.

When Claude, Gemini, OpenAI and Perplexity answered the same B2B buyer questions, they generally cited different exact pages. Across the six provider pairs, average within-question URL Jaccard overlap ranged from 0.93% to 2.02%. Domain and publisher overlap were slightly higher, but both remained below 5% for every pair.

  • Central finding

    The same buyer question did not lead the four provider stacks to converge on one common evidence set. Each stack frequently assembled its answer from a different part of the available information environment.

This does not mean 98–99% of every answer was different. It means the observed cited URL sets had very little intersection relative to their combined size. Nor does it mean the systems used completely different sources: the measured intersections were low, not zero.

Finding 1

The exact-URL overlap was extremely low

Exact URL overlap is the strictest comparison: two providers only match when they cite the same canonical page for the same buyer question. All six provider pairs remained at or below approximately 2.02% average URL Jaccard overlap.

Average within-question exact-URL Jaccard overlap by provider pair
Provider pair URL Jaccard overlap
Claude–OpenAI2.02%
Claude–Gemini2.01%
Gemini–OpenAI1.70%
Gemini–Perplexity1.52%
OpenAI–Perplexity1.01%
Claude–Perplexity0.93%
Provider-pair Jaccard overlap for cited URLs, domains and publishers across Claude, Gemini, OpenAI and Perplexity, with very low overlap throughout.
Figure 1. Average within-question Jaccard overlap across the nine-question matched panel. URL, domain and publisher overlap are separate measurements. Low overlap in this fixed benchmark does not establish repeated-run stability.
Open full-resolution figure

The differences between individual pairs are less important than the shared pattern: every pair had very low exact-page convergence. A citation observed in one provider's evidence set should not be assumed to imply exposure in another.

The study does not identify why a particular page was selected by one stack and not another. The observable result is narrower: the final cited evidence sets differed.

Finding 2

The pattern remains at domain and publisher level

Exact-page matching is demanding, so the study also compared the evidence at broader domain and publisher levels. Different URLs from the same site count as a domain match, while publisher resolution accounts for cases where a domain and publishing organisation are not identical.

Domain overlap

Average within-question domain Jaccard overlap by provider pair
Provider pair Domain Jaccard overlap
Claude–OpenAI4.76%
Gemini–Perplexity4.39%
Claude–Gemini3.89%
OpenAI–Perplexity3.27%
Gemini–OpenAI3.14%
Claude–Perplexity2.47%

The broader unit increases measured convergence, as expected, but not enough to create anything resembling a shared source universe.

Publisher overlap

Average within-question publisher Jaccard overlap by provider pair
Provider pair Publisher Jaccard overlap
Gemini–Perplexity4.79%
Claude–OpenAI4.76%
Claude–Gemini4.26%
Gemini–OpenAI3.49%
OpenAI–Perplexity3.27%
Claude–Perplexity1.23%

Cross-provider convergence remained low at every headline-eligible entity level. The provider stacks often reached into different domains and publishing organisations, not merely different pages from the same small publisher set.

Measure

What Jaccard overlap means

Jaccard overlap compares the intersection of two cited-source sets with their union. It answers: how similar are these evidence sets relative to the total evidence they collectively used?

Definition
Jaccard = shared sources ÷ all unique sources appearing in either set.
Example
If two providers each cite five URLs and share one, the union is nine URLs and Jaccard is 1 ÷ 9 = 11.1%.
Study calculation
Jaccard is calculated within each matched buyer question and then summarised across the fixed question panel. It is not one pooled overlap after combining every citation.

Sensitivity analysis

A second overlap measure tells the same broad story

Jaccard can become smaller when one provider cites substantially more sources than another because the union grows. The overlap coefficient instead divides shared sources by the size of the smaller set. It asks how much of the smaller evidence set was also present in the larger set.

Average within-question exact-URL overlap coefficient by provider pair
Provider pair URL overlap coefficient
Claude–Gemini6.10%
Gemini–Perplexity4.85%
Gemini–OpenAI4.44%
Claude–OpenAI3.70%
Claude–Perplexity1.85%
OpenAI–Perplexity1.85%
Horizontal dot plot of average URL overlap coefficient for all six provider pairs, ranging from 1.85% to 6.10%.
Figure 2. Average within-question URL overlap coefficient across the matched panel. The smaller source set is the denominator. This remains a descriptive fixed-panel measure, not an estimate of repeated-run stability or provider-wide source preference.
Open full-resolution figure

The values are larger than the corresponding Jaccard values, as expected, but exact-URL overlap still ranged only from approximately 1.85% to 6.10%. The low Jaccard values are therefore not only a consequence of the union denominator.

Context

Why source-set size matters

The four provider stacks did not cite the same number of domains. Across the nine-question panel, their average unique cited-domain breadth was:

Average unique cited domains per matched buyer question
Provider Unique cited domains per question
Gemini9.22
OpenAI6.22
Perplexity5.89
Claude5.33

A broader evidence set creates more opportunities for both matching and non-matching. That is why the study reports several entity levels and both Jaccard and overlap coefficient instead of reducing evidence convergence to one number.

Overlap measures convergence—not source authority, factual accuracy or answer quality. High convergence could reflect shared authoritative evidence or a narrow source ecosystem. Low convergence could reflect useful complementary evidence, fragmentation or instability. The experiment does not determine whether convergence itself is desirable.

Practical interpretation

What this means for AI Answer Alignment

AI visibility should not be treated as a single-system property. One provider may cite a company's authoritative explanation; another may cite a competitor, agency, directory or editorial source. Monitoring final answers without inspecting their evidence can hide those differences.

  • Representation

    Is the company described consistently across providers?

  • Supporting evidence

    Do the provider stacks rely on the same sources and proof points?

  • Third-party influence

    Which publishers, competitors, consultancies or reviewers help shape the answer?

  • Proof-point coverage

    Are important facts available across the wider public information landscape?

  • Narrative risk

    Are inaccurate or outdated claims reinforced elsewhere?

This is why Kojable treats AI Answer Alignment as broader than mention tracking. The question is not only whether an AI mentions a company, but what information environment produced that representation and whether it contains the evidence needed to tell the correct story.

Interpretation boundary

The provider pairs should not be ranked from this chart

Pairwise values make it tempting to say that one pair agrees more than another. For example, Claude–OpenAI has the highest URL Jaccard point estimate at 2.02%, while Claude–Perplexity has the lowest at 0.93%.

These are averages from nine matched questions with one observed response per provider-question cell. The study was not designed to establish a stable provider-pair hierarchy. The defensible finding is the common pattern: all six provider pairs showed low cross-provider source convergence.

Evidence boundaries

What this study does not establish

  • Answer agreement

    Two providers can cite different evidence and reach similar conclusions, or cite overlapping evidence and interpret it differently. Answer and evidence similarity are separate.

  • Source quality

    Shared sources are not necessarily authoritative and unique sources are not necessarily weak. Overlap does not measure factual accuracy or answer quality.

  • Run-to-run stability

    One observed run per provider-question cell cannot establish whether a provider would cite the same evidence in another session, generation or collection period.

  • Causal mechanism

    The observed stacks can combine generated searches, search backends, retrieval, ranking, orchestration, source metadata, citation implementation and answer-generation behaviour. The study does not isolate which mechanism caused divergence.

  • Universal generalisation

    Ten designed questions in one B2B AI-visibility and AI-answer-alignment context are a defined benchmark, not a random sample of every buyer question, industry or provider use case.

Method

Study design

This companion uses the same dataset as the flagship Cross-Model Citation Study. It is not a separate experiment or independent replication.

Designed panel
10 B2B buyer questions across four provider stacks, creating 40 expected provider-question responses.
Successful responses
39 responses succeeded. One Gemini response was missing.
Primary matched panel
Nine questions successfully completed by Claude, Gemini, OpenAI and Perplexity.
Analytical unit
Pairwise overlap is calculated within matched questions and summarised across the fixed panel. Each provider-question cell contains one observed run.
Headline entity levels
Exact canonical URL, domain and publisher pass the study's reporting-eligibility rules.
Named-author gate
Named-author overlap was not headline eligible: 0 of 9 matched questions met the author-coverage gate for every provider pair.

All six provider pairs

  • Claude–Gemini
  • Claude–OpenAI
  • Claude–Perplexity
  • Gemini–OpenAI
  • Gemini–Perplexity
  • OpenAI–Perplexity

Named-author overlap is supplementary and intentionally excluded from the main figures and headline findings in this companion.

Research record

Research and reproducibility

This companion preserves the same provider-question responses, canonicalisation rules, nine-question complete-case panel, provider-pair definitions, entity-level definitions and benchmark-panel uncertainty framework as the flagship study.

The overlap values come from the study's locked publication outputs. No new analytical metric was introduced for this article.

The underlying publication package is retained, but public repository access is currently unavailable. Until that route is restored, this page does not present the complete package as publicly accessible.

Hero
Reused editorial presentation of the locked exact-URL Jaccard range.
Figure 1
Reused flagship comparison of URL, domain and publisher Jaccard overlap for all six provider pairs.
Figure 2
New web presentation derived from the locked exact-URL overlap coefficient rows for all six provider pairs.

FAQ

Frequently asked questions

Do Claude, Gemini, OpenAI and Perplexity use the same sources?

They showed very little exact-source convergence in this fixed benchmark. Average within-question URL Jaccard overlap across the six provider pairs ranged from approximately 0.9% to 2.0%.

Which two AI systems had the most similar citations?

Claude and OpenAI had the highest URL Jaccard point estimate at approximately 2.02%, narrowly above Claude and Gemini at 2.01%. With only nine matched questions and one observed run per cell, this should not be interpreted as a stable provider-pair ranking.

Was domain overlap higher than exact URL overlap?

Yes. Domain Jaccard overlap ranged from approximately 2.47% to 4.76%, compared with approximately 0.93% to 2.02% for exact URLs.

Was publisher overlap higher?

Publisher Jaccard overlap ranged from approximately 1.23% to 4.79%. As with domain overlap, it remained low across all six provider pairs.

What is the difference between Jaccard overlap and overlap coefficient?

Jaccard divides the shared sources by the union of the two source sets. The overlap coefficient divides the shared sources by the size of the smaller source set. The second measure is less penalised by unequal set sizes. In this benchmark, URL overlap remained low under both measures.

Does low source overlap mean the AI answers were inaccurate?

No. This study measures whether providers cited the same evidence, not whether the final answers were factually correct. Different evidence sets can support similar conclusions.

Does this prove each provider has a unique source preference?

No. Each provider-question cell contains one observed run. Repeated sessions are needed to separate persistent provider-stack differences from normal run-to-run variability.

From benchmark to company evidence

See what this looks like for your company

Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.

See how Kojable works →
Explore AI citation monitoring →

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish