Kojable research · Cross-Model Citation Study

The Source Ecosystems Behind Claude, Gemini, OpenAI and Perplexity

Published By Piush Vaish

The same B2B buyer questions led four AI provider stacks to cite noticeably different mixtures of independent publishers, commercially interested sources, competitor pages and first-party material.

Key finding

Observed independent-source share ranged from 22.0% to 65.1%, commercially interested-source share from 34.9% to 83.6%, and direct-competitor share from 7.8% to 38.7%.

Protocol qualification: The responses came from an instructed research protocol that asked for criteria including independent, practitioner-led, non-Kojable and non-competitor sources, with -site:kojable.com in shared queries. The observed mix reflects both provider-stack behaviour and a research brief that shaped retrieval—not organic provider preferences in unconstrained use.

  • 4 provider stacks
  • 9 matched questions
  • 281 cited canonical URLs
  • 11 min read
Separate comparisons of independent-source share and commercial-interest share across Claude, Gemini, OpenAI and Perplexity under the instructed research protocol.
Independent ownership and commercial interest are separate, overlapping taxonomy dimensions. Source relationship is descriptive, not a quality score.
Open full-resolution figure
4provider stacks
9questions in the matched panel
281observed cited canonical URLs
100%cited-source taxonomy coverageInterpretationResults reflect an instructed research protocol.

Answer first

Did the four systems cite the same kinds of sources?

Direct answer

The four provider stacks cited materially different mixtures of source relationships in this benchmark.

The values are question-macro estimates from the nine-question matched panel. They describe overlapping taxonomy dimensions rather than mutually exclusive parts of a whole, and they are not quality scores.

Independent, commercially interested, direct-competitor and Kojable first-party source shares shown separately for Claude, Gemini, OpenAI and Perplexity.
Figure 1. Observed source-relationship shares in the nine-question matched panel. The dimensions overlap and are not stacked. “Independent” is a taxonomy classification rather than a judgement of authority, factual accuracy or objectivity.
Open full-resolution figure
Source-relationship shares in the nine-question matched panel
ProviderIndependentCommercial interestDirect competitorKojable first party
Claude65.1%34.9%7.8%0.0%
Gemini34.2%70.5%24.2%7.1%
OpenAI44.4%57.4%22.1%3.6%
Perplexity22.0%83.6%38.7%0.0%

Taxonomy boundary

Independent and commercial are not simple opposites

The taxonomy separates three analytical dimensions to avoid turning unlike properties into one source label.

Source format
Article/editorial, news, documentation, product page, comparison listicle, review directory, research report, academic paper, community post, video or case study.
Ownership type
Kojable first party, direct competitor, adjacent vendor, agency or consultancy, independent publisher, academic or institutional, or community/user-generated.
Commercial interest
Direct, indirect or adjacent commercial interest; affiliate or directory interest; independent editorial; public or academic; or community non-commercial or mixed.

A source can be editorial in format, independently owned and still have some market relationship depending on the dimension being discussed. A consultancy article, for example, can be editorial, commercially interested and neither first party nor a direct competitor.

The label describes the relationship and form of the evidence. It does not automatically describe its quality.

Finding 1

The provider stacks produced different source mixes

Independent-source share with benchmark-panel intervals
ProviderIndependent-source shareBenchmark-panel interval
Claude65.1%45.0%–82.9%
OpenAI44.4%32.6%–58.1%
Gemini34.2%26.0%–43.8%
Perplexity22.0%4.8%–40.5%
Commercially interested-source share with benchmark-panel intervals
ProviderCommercial-interest shareBenchmark-panel interval
Perplexity83.6%68.4%–98.0%
Gemini70.5%61.9%–79.0%
OpenAI57.4%47.0%–67.5%
Claude34.9%17.1%–55.0%

The largest contrast was Claude versus Perplexity: approximately 43 percentage points for independent-source share and 49 percentage points for commercial-interest share. Under the study's existing matched-question uncertainty and multiplicity procedure, these contrasts remained robust to the composition of the fixed question panel.

Most other provider-pair differences should remain descriptive. This is not a league table, a permanent provider ranking or evidence that one source mix is inherently better.

Finding 2

Competitor sources were unevenly represented

Direct-competitor source share in the matched panel
ProviderDirect-competitor share
Perplexity38.7%
Gemini24.2%
OpenAI22.1%
Claude7.8%

Competitor content can define category language, evaluation criteria, comparison framing and perceived trade-offs. That does not make it inherently inappropriate: a competitor's framework may be directly relevant to a buyer's question.

The representation risk appears when a company's category story exists primarily through competitors rather than through a broader, well-supported public information environment.

Finding 3

First-party citations were sparse and protocol-sensitive

Observed Kojable first-party source share
ProviderKojable first-party share
Claude0.0%
Gemini7.1%
OpenAI3.6%
Perplexity0.0%

The underlying workflow requested non-Kojable evidence and used -site:kojable.com in shared search strings. It would therefore be inappropriate to infer that Claude or Perplexity generally do not use first-party sources.

First-party evidence is often strongest for official product facts, pricing, documentation, policy and release information. Third-party sources play a different role in validation, comparison, reputation and market context. A robust information environment normally needs both.

Finding 4

Most cited pages were editorial articles

Article/editorial was the dominant observed format for all four provider stacks, despite their differing ownership and commercial-interest mixes.

Provider-by-format heatmap showing question-macro cited-source shares across the locked source-format categories for Claude, Gemini, OpenAI and Perplexity.
Figure 2. Question-macro source-format share by provider. The locked taxonomy is unchanged; zero-only display rows may be omitted for legibility but remain in the underlying table. Format is descriptive and does not establish that a page type causes AI citation.
Open full-resolution figure
Article/editorial page-format share in the matched panel
ProviderArticle/editorial share
Claude79.0%
Perplexity68.5%
OpenAI63.7%
Gemini60.8%

This observational result does not mean companies should publish editorial articles because AI systems “prefer” them. The instructed protocol, available content, retrieval systems and selection process can all affect the observed format mix. A causal recommendation requires a separate intervention study.

Method

How the source taxonomy works

The study used source-taxonomy-v1.0.0. The classifier is deterministic and outcome-blind: one canonical URL is classified once in the source master, then reused downstream.

Source master
1,637 canonical URLs across the full successful-response set.
Observed cited set
281 cited canonical URLs across 39 successful provider-question responses.
Coverage
100% for cited-source format, ownership and commercial-interest classifications. Coverage is not the same as classification accuracy.
  • Classifier inputs

    Canonical URL, hostname, registrable domain, source title and provider source-type label.

  • Withheld outcomes

    The classifier does not receive citation status, retrieval status, citation-event count or candidate-selection outcome.

  • other and unknown

    other is classified but outside more specific values; unknown means the evidence is insufficient to classify the dimension.

Interpretation

Why source type is not source quality

The analysis describes what kind of source appeared. It does not determine whether that source was factually correct, authoritative, current, unbiased, comprehensive or the best evidence for the claim.

  • Independent does not mean authoritative

    An independently owned publisher can still be wrong, outdated or poorly sourced.

  • Commercial does not mean weak

    Official documentation or a product page may be definitive evidence for features, specifications or policies.

  • Competitor does not mean irrelevant

    A competitor may publish detailed educational or comparison material that directly addresses a buyer question.

The goal is not to maximise independent-source share. It is to ensure that the evidence ecosystem contains accurate, appropriate and corroboratable information for the claims buyers care about.

Company implications

What this means for companies

The source ecosystem around a company is broader than its own website. AI systems can construct buyer-facing answers from first-party pages, competitors, agencies, consultants, independent publishers, research organisations, directories, communities and market commentary.

  1. First-party content is necessary but not sufficient

    A company can publish correct information and still be represented through third-party sources.

  2. Competitor framing can become evidence

    Clearer category explanations and comparison material can shape the public information environment around the market.

  3. Independent corroboration plays a different role

    Official facts may belong on first-party pages, while reputation, outcomes and market position benefit from external support.

  4. Diagnose source mix claim by claim

    Ask whether the right source type supports the right claim rather than treating one taxonomy category as universally preferable.

The practical chain is buyer question → AI answer → claim → cited source → source relationship → support → representation.

Study design

A fixed matched-question benchmark

This companion uses the same dataset as the flagship Cross-Model Citation Study: 10 designed B2B buyer questions, four provider stacks, 40 expected provider-question cells, 39 successful responses, one missing Gemini response and nine questions in the primary four-provider matched panel.

Each provider-question cell contains one observed run. Provider values are question-macro estimates. Benchmark-panel intervals describe sensitivity to the composition of this fixed panel, not run-to-run stability or universal provider traits.

Evidence boundaries

Limitations

  • Instructed research protocol

    Shared queries asked for independent, practitioner-led, non-Kojable and non-competitor evidence and included -site:kojable.com. The observed mix is protocol-shaped.

  • Market-relative taxonomy

    Commercial interest is defined relative to the market context, not as a universal property of a domain.

  • Classification is not quality review

    Format, ownership and commercial-interest categories do not assess factual support, authority or correctness.

  • Supplementary measures

    Competitor and first-party shares are useful diagnostics but are not all primary headline outcomes under the locked scorecard.

  • Final cited sources

    Candidate pools are not equally observable across all four providers, so this companion compares final cited-source composition.

  • No causal content rule

    Editorial-format prevalence does not prove that publishing an article causes citation. Repeated intervention experiments are required.

Research record

Research and reproducibility

The public package contains the source-taxonomy methodology, source-ecosystem tables, classification coverage, provider scorecards, benchmark-panel intervals, source-master reconciliation and publication manifests.

View the Cross-Model Citation Study publication package on GitHub.

FAQ

Frequently asked questions

Which AI provider cited the highest share of independent sources?

Claude had the highest observed independent-source share in the nine-question matched panel at 65.1%. OpenAI was 44.4%, Gemini 34.2% and Perplexity 22.0%.

Which provider cited the most commercially interested sources?

Perplexity had the highest observed commercially interested-source share at 83.6%, followed by Gemini at 70.5%, OpenAI at 57.4% and Claude at 34.9%.

Are independent and commercial sources mutually exclusive?

No. The study separates ownership from commercial-interest classification. The dimensions can overlap and should not be stacked into a single 100% composition.

Does a commercial source mean a low-quality source?

No. Commercial interest describes the source's relationship to the market. Official product documentation, for example, can be commercially interested and still be the best source for a product fact.

Does an independent source mean an authoritative source?

No. Independence is not a factual-accuracy or authority score.

Which provider cited the highest share of direct competitors?

Perplexity had the highest observed direct-competitor share at 38.7% in the matched panel. Gemini was 24.2%, OpenAI 22.1% and Claude 7.8%. This is a supplementary taxonomy measure.

Why were first-party Kojable citations so low?

The instructed research workflow explicitly included search criteria designed to find non-Kojable evidence and used -site:kojable.com in shared queries. First-party results are therefore especially protocol-sensitive and should not be interpreted as general provider behaviour.

What page format was cited most often?

Article/editorial pages were the largest observed format category for all four provider stacks, ranging from approximately 60.8% for Gemini to 79.0% for Claude.

Does that mean companies should publish more editorial articles to get cited?

This study cannot establish that. The format result is observational and protocol-dependent. A causal recommendation would require an intervention experiment.

From benchmark to company evidence

See what this looks like for your company

Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.

See how Kojable works →
Explore AI citation monitoring →

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish