Kojable research · Cross-Model Citation Study
The Source Ecosystems Behind Claude, Gemini, OpenAI and Perplexity
The same B2B buyer questions led four AI provider stacks to cite noticeably different mixtures of independent publishers, commercially interested sources, competitor pages and first-party material.
Key finding
Observed independent-source share ranged from 22.0% to 65.1%, commercially interested-source share from 34.9% to 83.6%, and direct-competitor share from 7.8% to 38.7%.
Protocol qualification: The responses came from an instructed research protocol that asked for criteria including independent, practitioner-led, non-Kojable and non-competitor sources, with -site:kojable.com in shared queries. The observed mix reflects both provider-stack behaviour and a research brief that shaped retrieval—not organic provider preferences in unconstrained use.
- 4 provider stacks
- 9 matched questions
- 281 cited canonical URLs
- 11 min read
Answer first
Did the four systems cite the same kinds of sources?
Direct answer
The four provider stacks cited materially different mixtures of source relationships in this benchmark.
The values are question-macro estimates from the nine-question matched panel. They describe overlapping taxonomy dimensions rather than mutually exclusive parts of a whole, and they are not quality scores.

| Provider | Independent | Commercial interest | Direct competitor | Kojable first party |
|---|---|---|---|---|
| Claude | 65.1% | 34.9% | 7.8% | 0.0% |
| Gemini | 34.2% | 70.5% | 24.2% | 7.1% |
| OpenAI | 44.4% | 57.4% | 22.1% | 3.6% |
| Perplexity | 22.0% | 83.6% | 38.7% | 0.0% |
Taxonomy boundary
Independent and commercial are not simple opposites
The taxonomy separates three analytical dimensions to avoid turning unlike properties into one source label.
- Source format
- Article/editorial, news, documentation, product page, comparison listicle, review directory, research report, academic paper, community post, video or case study.
- Ownership type
- Kojable first party, direct competitor, adjacent vendor, agency or consultancy, independent publisher, academic or institutional, or community/user-generated.
- Commercial interest
- Direct, indirect or adjacent commercial interest; affiliate or directory interest; independent editorial; public or academic; or community non-commercial or mixed.
A source can be editorial in format, independently owned and still have some market relationship depending on the dimension being discussed. A consultancy article, for example, can be editorial, commercially interested and neither first party nor a direct competitor.
The label describes the relationship and form of the evidence. It does not automatically describe its quality.
Finding 1
The provider stacks produced different source mixes
| Provider | Independent-source share | Benchmark-panel interval |
|---|---|---|
| Claude | 65.1% | 45.0%–82.9% |
| OpenAI | 44.4% | 32.6%–58.1% |
| Gemini | 34.2% | 26.0%–43.8% |
| Perplexity | 22.0% | 4.8%–40.5% |
| Provider | Commercial-interest share | Benchmark-panel interval |
|---|---|---|
| Perplexity | 83.6% | 68.4%–98.0% |
| Gemini | 70.5% | 61.9%–79.0% |
| OpenAI | 57.4% | 47.0%–67.5% |
| Claude | 34.9% | 17.1%–55.0% |
The largest contrast was Claude versus Perplexity: approximately 43 percentage points for independent-source share and 49 percentage points for commercial-interest share. Under the study's existing matched-question uncertainty and multiplicity procedure, these contrasts remained robust to the composition of the fixed question panel.
Most other provider-pair differences should remain descriptive. This is not a league table, a permanent provider ranking or evidence that one source mix is inherently better.
Finding 2
Competitor sources were unevenly represented
| Provider | Direct-competitor share |
|---|---|
| Perplexity | 38.7% |
| Gemini | 24.2% |
| OpenAI | 22.1% |
| Claude | 7.8% |
Competitor content can define category language, evaluation criteria, comparison framing and perceived trade-offs. That does not make it inherently inappropriate: a competitor's framework may be directly relevant to a buyer's question.
The representation risk appears when a company's category story exists primarily through competitors rather than through a broader, well-supported public information environment.
Finding 3
First-party citations were sparse and protocol-sensitive
| Provider | Kojable first-party share |
|---|---|
| Claude | 0.0% |
| Gemini | 7.1% |
| OpenAI | 3.6% |
| Perplexity | 0.0% |
The underlying workflow requested non-Kojable evidence and used -site:kojable.com in shared search strings. It would therefore be inappropriate to infer that Claude or Perplexity generally do not use first-party sources.
First-party evidence is often strongest for official product facts, pricing, documentation, policy and release information. Third-party sources play a different role in validation, comparison, reputation and market context. A robust information environment normally needs both.
Finding 4
Most cited pages were editorial articles
Article/editorial was the dominant observed format for all four provider stacks, despite their differing ownership and commercial-interest mixes.

| Provider | Article/editorial share |
|---|---|
| Claude | 79.0% |
| Perplexity | 68.5% |
| OpenAI | 63.7% |
| Gemini | 60.8% |
This observational result does not mean companies should publish editorial articles because AI systems “prefer” them. The instructed protocol, available content, retrieval systems and selection process can all affect the observed format mix. A causal recommendation requires a separate intervention study.
Method
How the source taxonomy works
The study used source-taxonomy-v1.0.0. The classifier is deterministic and outcome-blind: one canonical URL is classified once in the source master, then reused downstream.
- Source master
- 1,637 canonical URLs across the full successful-response set.
- Observed cited set
- 281 cited canonical URLs across 39 successful provider-question responses.
- Coverage
- 100% for cited-source format, ownership and commercial-interest classifications. Coverage is not the same as classification accuracy.
- Classifier inputs
Canonical URL, hostname, registrable domain, source title and provider source-type label.
- Withheld outcomes
The classifier does not receive citation status, retrieval status, citation-event count or candidate-selection outcome.
otherandunknownotheris classified but outside more specific values;unknownmeans the evidence is insufficient to classify the dimension.
Interpretation
Why source type is not source quality
The analysis describes what kind of source appeared. It does not determine whether that source was factually correct, authoritative, current, unbiased, comprehensive or the best evidence for the claim.
- Independent does not mean authoritative
An independently owned publisher can still be wrong, outdated or poorly sourced.
- Commercial does not mean weak
Official documentation or a product page may be definitive evidence for features, specifications or policies.
- Competitor does not mean irrelevant
A competitor may publish detailed educational or comparison material that directly addresses a buyer question.
The goal is not to maximise independent-source share. It is to ensure that the evidence ecosystem contains accurate, appropriate and corroboratable information for the claims buyers care about.
Company implications
What this means for companies
The source ecosystem around a company is broader than its own website. AI systems can construct buyer-facing answers from first-party pages, competitors, agencies, consultants, independent publishers, research organisations, directories, communities and market commentary.
- First-party content is necessary but not sufficient
A company can publish correct information and still be represented through third-party sources.
- Competitor framing can become evidence
Clearer category explanations and comparison material can shape the public information environment around the market.
- Independent corroboration plays a different role
Official facts may belong on first-party pages, while reputation, outcomes and market position benefit from external support.
- Diagnose source mix claim by claim
Ask whether the right source type supports the right claim rather than treating one taxonomy category as universally preferable.
The practical chain is buyer question → AI answer → claim → cited source → source relationship → support → representation.
Study design
A fixed matched-question benchmark
This companion uses the same dataset as the flagship Cross-Model Citation Study: 10 designed B2B buyer questions, four provider stacks, 40 expected provider-question cells, 39 successful responses, one missing Gemini response and nine questions in the primary four-provider matched panel.
Each provider-question cell contains one observed run. Provider values are question-macro estimates. Benchmark-panel intervals describe sensitivity to the composition of this fixed panel, not run-to-run stability or universal provider traits.
Evidence boundaries
Limitations
- Instructed research protocol
Shared queries asked for independent, practitioner-led, non-Kojable and non-competitor evidence and included
-site:kojable.com. The observed mix is protocol-shaped. - Market-relative taxonomy
Commercial interest is defined relative to the market context, not as a universal property of a domain.
- Classification is not quality review
Format, ownership and commercial-interest categories do not assess factual support, authority or correctness.
- Supplementary measures
Competitor and first-party shares are useful diagnostics but are not all primary headline outcomes under the locked scorecard.
- Final cited sources
Candidate pools are not equally observable across all four providers, so this companion compares final cited-source composition.
- No causal content rule
Editorial-format prevalence does not prove that publishing an article causes citation. Repeated intervention experiments are required.
Research record
Research and reproducibility
The public package contains the source-taxonomy methodology, source-ecosystem tables, classification coverage, provider scorecards, benchmark-panel intervals, source-master reconciliation and publication manifests.
View the Cross-Model Citation Study publication package on GitHub.
FAQ
Frequently asked questions
Which AI provider cited the highest share of independent sources?
Claude had the highest observed independent-source share in the nine-question matched panel at 65.1%. OpenAI was 44.4%, Gemini 34.2% and Perplexity 22.0%.
Which provider cited the most commercially interested sources?
Perplexity had the highest observed commercially interested-source share at 83.6%, followed by Gemini at 70.5%, OpenAI at 57.4% and Claude at 34.9%.
Are independent and commercial sources mutually exclusive?
No. The study separates ownership from commercial-interest classification. The dimensions can overlap and should not be stacked into a single 100% composition.
Does a commercial source mean a low-quality source?
No. Commercial interest describes the source's relationship to the market. Official product documentation, for example, can be commercially interested and still be the best source for a product fact.
Does an independent source mean an authoritative source?
No. Independence is not a factual-accuracy or authority score.
Which provider cited the highest share of direct competitors?
Perplexity had the highest observed direct-competitor share at 38.7% in the matched panel. Gemini was 24.2%, OpenAI 22.1% and Claude 7.8%. This is a supplementary taxonomy measure.
Why were first-party Kojable citations so low?
The instructed research workflow explicitly included search criteria designed to find non-Kojable evidence and used -site:kojable.com in shared queries. First-party results are therefore especially protocol-sensitive and should not be interpreted as general provider behaviour.
What page format was cited most often?
Article/editorial pages were the largest observed format category for all four provider stacks, ranging from approximately 60.8% for Gemini to 79.0% for Claude.
Does that mean companies should publish more editorial articles to get cited?
This study cannot establish that. The format result is observational and protocol-dependent. A causal recommendation would require an intervention experiment.
From benchmark to company evidence
See what this looks like for your company
Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.
