Kojable research · Cross-Model Citation Study
Different Answers, Different Evidence: How Claude, Gemini, OpenAI and Perplexity Cite the Web
The same buyer question can produce four plausible AI answers—and four largely different evidence sets.
Key finding
The four provider stacks frequently assembled materially different evidence environments around the same buyer questions. Average pairwise URL Jaccard overlap was only about 0.009–0.020, while visible citation density ranged from 5.28 citation events per 1,000 words for OpenAI to 39.79 for Gemini.
Qualification: This is a fixed observational benchmark with one observed run per provider-question cell. It does not identify a universal model preference, provider-wide accuracy ranking, causal citation mechanism or run-to-run stability.
- Cross-provider AI citations
- 10 designed questions
- 39 successful responses
- 14 min read
Study overview
Executive summary
AI answers can look similar on the surface while being built from very different evidence underneath. Kojable presented the same ten B2B buyer questions to Claude, Gemini, OpenAI and Perplexity and analysed the citations, source mix, freshness, support relationships, publishers and authors behind the responses.
Of 40 expected provider-question cells, 39 succeeded. Gemini P08 returned no usable answer, so the primary four-provider comparison uses the nine questions completed by all four provider stacks.
-
Citation density
Ranged from 5.28 citation events per 1,000 words for OpenAI to 39.79 for Gemini.
-
Cited-domain breadth
Averaged 5.33 unique domains per question for Claude and 9.22 for Gemini at the observed range endpoints.
-
Source mix
Observed independent-source share ranged from 22.0% for Perplexity to 65.1% for Claude.
-
Freshness
Median cited-source age ranged from approximately 76 days for Claude to 120 days for Perplexity, conditional on publication-date coverage.
-
Citation placement
Automated material-claim citation coverage ranged from 10.2% for OpenAI to 43.2% for Gemini.
-
Reviewed support
Direct-support rates among reviewed relationships ranged from 72.0% for Gemini to 86.7% for OpenAI, but provider review coverage differed substantially.
-
Evidence overlap
Pairwise URL overlap was extremely low: average within-question Jaccard values ranged from roughly 0.009 to 0.020.
-
Named authors
Named-author share ranged from 50.9% for Gemini to 77.1% for Claude.
Answer first
The four provider stacks did not rely on the same evidence
Direct answer
Claude, Gemini, OpenAI and Perplexity differed substantially in visible citation density, cited-domain breadth, source mix and the individual URLs and publishers that appeared. The clearest signal was low source overlap: the same question often led the provider stacks to assemble largely different cited-source sets.
Representation in one provider's evidence environment cannot be assumed to transfer to another. Monitoring only one AI system can therefore miss part of the public information landscape associated with AI-mediated buyer research.
The study does not show that one provider is universally better, more accurate or more authoritative. It also does not establish that the observed source differences are stable across repeated runs. For definitions separating citations, retrieval, grounding and source influence, read Kojable's AI citation guide.
Finding 1
Citation intensity varied dramatically
Gemini produced the highest visible citation density in the matched panel at 39.79 citation events per 1,000 words. Perplexity followed at 21.06, Claude at 8.40, and OpenAI at 5.28.
A citation event records that a source was attached to an answer placement. It does not establish that the source is authoritative, supports the claim correctly or makes the answer factually better. Nor does a larger number of visible citations necessarily mean that a provider consulted more evidence internally.
The defensible conclusion is narrower: under the same benchmark protocol, the provider stacks produced very different levels of visible citation activity. Raw citation counts should not be treated as directly interchangeable across systems.
Finding 2
Evidence breadth and concentration also differed
Gemini showed the broadest cited-domain profile, averaging 9.22 unique cited domains per question. OpenAI averaged 6.22, Perplexity 5.89, and Claude 5.33. Gemini and OpenAI had lower observed domain HHI values, while Claude and Perplexity were more concentrated.
A broader set can expose a reader to more perspectives, but breadth is not a quality score. A concentrated answer may rely on a small number of strong sources; a broad answer may contain weak, repetitive or commercially similar sources. Breadth and concentration therefore remain separate measurements.
Finding 3
The source ecosystems were materially different
Cited pages were classified across separate dimensions including ownership, commercial interest and content format. An independent source is a taxonomy category, not a guarantee of objectivity or factual quality. Commercially interested, competitor and first-party sources can all contain accurate and useful information.
Observed independent-source share ranged from 65.1% for Claude to 22.0% for Perplexity. Gemini was 34.2% and OpenAI 44.4%. The commercially interested-source dimension ranged from 34.9% for Claude to 83.6% for Perplexity.
The provider stacks drew on different mixtures of publishers, vendors, agencies, competitors, communities and first-party sources. AI representation is therefore not determined only by what exists on a company's own website.
- First-party product information
- Competitor narratives
- Agency or consultancy content
- Independent editorial coverage
- Review, directory and community pages
- Research or institutional material
Kojable's Monitor → Diagnose → Improve → Verify workflow separates the observed answer from the source and information gaps associated with it.
Finding 4
Cited-source freshness also differed
Source age was measured relative to the timestamp when each response was generated. Missing or ambiguous publication dates remain unknown rather than being assigned to an old-content bucket.
| Provider | Median cited-source age | Publication-date coverage |
|---|---|---|
| Claude | 76 days | 88.8% |
| Gemini | 81 days | 64.8% |
| OpenAI | 88 days | 70.3% |
| Perplexity | 120 days | 97.2% |
These results do not show that newer content causes citations, that older pages are penalised or that companies should publish on a particular cadence. Page characteristics such as word count, headings, images, lists, tables and structured-data markers are also descriptive; turning them into causal rules requires an intervention study.
Finding 5
Citation quantity, placement and support are different questions
- Citation intensity
- How many visible citation events appear.
- Material-claim citation coverage
- How often automatically identified material claims have an observed citation placement.
- Reviewed support
- Whether a cited source directly or partially supports the linked claim in the reviewed sample.
Automated material-claim citation coverage was 43.2% for Gemini, 33.7% for Perplexity, 25.0% for Claude and 10.2% for OpenAI. For support quality, the study reviewed a stratified sample of 120 claim-citation relationships from 1,298 mapped relationships. Each sampled relationship was independently assessed by one human reviewer and one AI-assisted web-research reviewer, with disagreements adjudicated.
| Provider | Direct-support rate | Annotation coverage |
|---|---|---|
| Claude | 78.5% | 19.5% |
| Gemini | 72.0% | 4.6% |
| OpenAI | 86.7% | 31.3% |
| Perplexity | 79.3% | 7.3% |
Reviewed support metrics are not estimates of overall answer accuracy, hallucination rate or provider-wide factual reliability. Raw agreement between the two assessors was 77.5%, Cohen's kappa was approximately 0.477, and 27 disagreements required adjudication. A high number of citation markers does not automatically mean stronger evidentiary support.
Finding 6
Cross-provider evidence overlap was very low
Across provider pairs, average within-question URL Jaccard overlap ranged from approximately 0.009 to 0.020. Domain overlap was also low, with average Jaccard values roughly between 0.025 and 0.048. Publisher overlap remained low as well.
Answer similarity and evidence similarity are not the same. Two answers may reach similar conclusions from different publishers, domains or commercial ecosystems; two systems may also cite some of the same domains while emphasising different pages and claims.
Visibility in one provider's evidence environment does not guarantee visibility in another. Within this fixed benchmark, cross-provider evidence convergence was low, but repeated-run work is required to test stability.
Finding 7
Publisher and author visibility added another layer
A domain is not the same thing as a publisher, and a publisher is not the same thing as an author. The study resolved these entities separately where the available evidence allowed it.
- Claude
- 77.1% named-author share
- OpenAI
- 69.3% named-author share
- Perplexity
- 65.5% named-author share
- Gemini
- 50.9% named-author share
The results do not prove that named authors are preferred or that bylines cause AI citation. They show authorship as another observable layer of the evidence environment. Unresolved identities remain unresolved rather than being guessed, and author-review provenance remains explicit.
Finding 8
Retrieval and citation are not the same stage
Claude and OpenAI exports expose candidate/search pools that the study treats as complete under its locked observability contract. This supports a distinction between sources observed in an exposed candidate pool and sources ultimately cited.
The same analysis is not valid for Gemini and Perplexity because their exposed source metadata is citation-biased rather than a validated complete search pool. Candidate-selection analysis is valid only where a sufficiently complete candidate/search pool was observable.
In the nine-question complete-case panel, the observed candidate-to-citation selection rate was approximately 11.0% for Claude and 6.2% for OpenAI.
The difference does not mean Claude is better at selection. A broader candidate pool can naturally produce a lower selection fraction, and the denominator is partly a property of retrieval and instrumentation.
-
Candidate observation
Was the source observed in an eligible exposed candidate pool?
-
Citation selection
If observed, was the source selected for citation?
-
Claim relationship
If cited, which claim or section did it support, and did the cited page support that claim?
-
Provider comparison
How did those observable states differ across eligible provider surfaces?
Practical interpretation
What these findings mean for companies
The study supports treating AI representation as a multi-provider evidence problem, not only as an answer-ranking or mention-count problem. A company can be represented differently because provider stacks can assemble information from different parts of the web.
-
Observed visibility
Which pages about the company appear, and which are selected into final answers across provider surfaces?
-
Narrative environment
Which third-party publishers, authors and competitor pages shape the category narrative?
-
Evidence gaps
Are important proof points absent, outdated claims still visible or accurate first-party facts difficult to corroborate?
-
Provider-specific patterns
Does a provider rely on a narrow domain set or draw from a materially different evidence mix?
The study does not provide a universal recipe for getting cited by AI. It supports a measurement principle: look beyond the mention and inspect the evidence environment around the answer. Kojable's AI citation monitoring applies that distinction to company-specific buyer questions.
Method
Study design
- Designed panel
- Ten B2B buyer questions across Claude, Gemini, OpenAI and Perplexity: 40 expected provider-question cells.
- Successful responses
- Thirty-nine cells succeeded. Gemini P08 returned no usable answer.
- Primary matched panel
- Nine questions completed by all providers: P01, P02, P03, P04, P05, P06, P07, P09 and P10.
- Analytical unit
- The question is primary. Citation events, URLs and source records are nested within provider-question responses. Each provider-question cell contains one observed run.
- Provider summaries
- Equal-weight question-level macro averages are used for primary provider summaries.
- Resampling
- Matched-question Poisson block resampling applies the same question-level weight across eligible providers within a replicate.
Provider observability
Provider-native source metadata does not expose the same object on every surface. The eligibility distinction is part of the analysis, not a missing-data convenience.
| Provider | Exposed candidate pool | Candidate-selection analysis |
|---|---|---|
| Claude | Treated as complete | Eligible |
| OpenAI | Treated as complete | Eligible |
| Gemini | Citation-biased | Not eligible |
| Perplexity | Citation-biased | Not eligible |
A source absent from Gemini or Perplexity's exposed metadata cannot be described as “not retrieved”; their underlying candidate pools are not fully observable in the current exports.
Evidence boundaries
Limitations
-
Designed question panel
The ten questions were designed for this study rather than randomly sampled from all possible B2B buyer questions.
-
One observed run per cell
The study does not estimate run-to-run volatility or longitudinal stability.
-
Instructed protocol
The research instruction may influence search behaviour, source selection, answer structure and citation behaviour.
-
Provider-stack comparison
Results can combine model behaviour, search backend, query generation, retrieval orchestration, API instrumentation, citation implementation and answer style. The study does not isolate those components causally.
-
Unequal observability
Candidate-to-citation selection is restricted to Claude and OpenAI. Gemini and Perplexity do not expose validated complete candidate pools under the study contract.
-
Coverage-qualified enrichment
Publication-date, page-feature, publisher and author resolution are incomplete to varying degrees. Missing values remain unknown rather than absent or zero.
-
Stratified support review
Support review covers a subset of mapped claim-citation relationships and uses one human plus one AI-assisted reviewer. It is not a provider-wide accuracy estimate.
Most importantly, this is one snapshot of probabilistic systems. Repeated-run experiments are required before treating the observed differences as stable provider traits.
Research record
Research and reproducibility
The publication package makes the reported numbers traceable through the locked analysis contract, provider-observability rules, metric definitions, taxonomy and enrichment methods, provider scorecards, publication tables, benchmark-panel resampling outputs, claim-level traceability, and input/output manifests and checksums.
The publication layer does not create a composite provider score, declare an overall winner or introduce a causal claim about why the provider stacks differed.
FAQ
Frequently asked questions
Do Claude, Gemini, OpenAI and Perplexity cite the same sources?
Not usually in this fixed benchmark. Average pairwise URL Jaccard overlap across the matched questions was approximately 0.009–0.020, indicating that the observed provider stacks frequently cited largely different URL sets for the same buyer questions.
Which provider produced the most citations?
Gemini had the highest visible citation density in the matched panel at 39.79 citation events per 1,000 words. Perplexity was 21.06, Claude 8.40 and OpenAI 5.28. Citation density is a placement metric, not a measure of answer quality or factual accuracy.
Does producing more citations mean an AI answer is better supported?
No. Citation volume, material-claim citation placement and reviewed claim-to-source support are different measurements. A high citation count does not by itself establish that the sources directly support the claims or that the overall answer is more accurate.
Which provider cited the highest share of independent sources?
Claude had the highest observed independent-source share in the matched panel at 65.1%. Perplexity had the lowest at 22.0%. “Independent” is a source-taxonomy category and is not automatically a judgment of authority, objectivity or accuracy.
Why is candidate-to-citation selection reported only for Claude and OpenAI?
Under the study's observability contract, Claude and OpenAI expose candidate/search pools that are treated as complete for this analysis. Gemini and Perplexity expose citation-biased source metadata, so an equivalent retrieval-to-citation selection rate would not be valid.
Does this study show stable provider preferences?
No. Each provider-question cell contains one observed run. The benchmark shows differences in the recorded provider stacks under the tested protocol, but repeated runs are needed before those differences can be treated as stable provider traits.
From benchmark to company evidence
See what this looks like for your company
Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.