Kojable research · Cross-Model Citation Study

From Search Results to AI Citations: What Claude and OpenAI Selected

Published By Piush Vaish

Appearing in an AI system's exposed search or candidate pool is not the same thing as being cited in the final answer.

Key finding

Across the nine-question matched panel, the question-macro candidate-to-citation selection rate was approximately 11.0% for Claude and 6.2% for OpenAI. Most observed candidates did not become final citations.

Qualification: Candidate-selection analysis is valid only for Claude and OpenAI under the locked observability contract. Gemini and Perplexity expose citation-biased metadata rather than comparable complete candidate pools. These rates are descriptive operational diagnostics—not provider-quality scores, causal probabilities or repeated-run estimates.

  • 2 eligible provider stacks
  • 9-question primary panel
  • 10-question sensitivity
  • 10 min read
Editorial funnel showing that only a minority of exposed Claude and OpenAI candidate sources became final citations, with question-macro rates of 11.0% and 6.2%.
Being found is not the same as being cited. The comparison is limited to the two provider stacks with complete exposed candidate pools.
Open full-resolution figure
2eligible provider stacks
9questions in the primary panel
11.0%Claude question-macro rate
6.2%OpenAI question-macro rateScopeComplete exposed candidate pools only.

Answer first

How often did exposed candidates become citations?

Direct answer

Once a source appeared in an eligible exposed candidate pool, only a minority of those observed candidates were ultimately selected for citation: approximately 11.0% for Claude and 6.2% for OpenAI in the nine-question question-macro comparison.

These are equal-weight question-macro estimates, not pooled raw-row ratios. They do not rank provider quality, and comparable rates are not reported for Gemini or Perplexity because their exposed source denominators are citation-biased.

Candidate-to-citation selection in eligible complete exposed pools
ProviderNine-question primaryBenchmark-panel intervalTen-question sensitivity
Claude10.99%7.7%–14.5%10.95%
OpenAI6.25%5.3%–7.5%6.46%

The locked estimates are 10.99208219% and 6.24642793% in the primary panel, with all-available sensitivities of 10.94550555% and 6.45511847%, respectively. Display values are rounded for readability.

Stage model

Retrieval exposure and citation are different states

The exposed metadata supports a stage model without claiming access to hidden model state.

  1. Not observed in the exposed candidate pool

    The source is absent from the provider-exposed search or candidate records associated with the response.

  2. Observed as a candidate

    The page enters the observable retrieval environment, but this does not show that it influenced the answer.

  3. Selected for citation

    The candidate appears in the final cited-source set and becomes visible answer evidence.

  4. Mapped to a claim

    A citation event can be connected to a claim or answer segment.

  5. Evaluated for support

    A separate evidence layer asks whether the cited source actually supports the linked claim.

The concise sequence is candidate exposure → citation selection → claim mapping → evidentiary support. This companion focuses on the observable transition from candidate exposure to final citation.

Finding 1

Claude and OpenAI selected a minority of observed candidates

Candidate-to-citation selection estimates and benchmark-panel intervals for Claude and OpenAI in the nine-question primary panel and ten-question sensitivity.
Figure 1. Question-macro candidate-selection estimates for Claude and OpenAI, the only provider stacks with complete exposed pools under the study contract. The intervals reflect fixed-panel composition, not run-to-run stability.
Open full-resolution figure

Claude's point estimate was higher in this observed panel, but the central finding is not a provider winner. Candidate exposure and citation selection were materially different stages in both eligible stacks.

Estimand

Why the primary rate is question-macro

The headline calculation first measures selection inside each eligible provider-question cell and then averages questions with equal weight. This prevents a question with a large exposed pool from automatically dominating the provider summary.

Question-macro result
Answers: for a typical question in this fixed benchmark, what share of its exposed candidate pool became cited sources?
Pooled row ratio
Answers: what fraction of all candidate rows across all questions became cited rows? Questions with larger pools receive more weight.
Primary analytical unit
The question. Provider-question responses, candidate sources, cited sources and citation events are nested objects.

For example, selection rates of 20% from 20 candidates and 5% from 200 candidates produce a 12.5% equal-weight macro average but a 6.4% pooled ratio. Neither calculation is inherently wrong; they answer different questions. This study specifies the question-macro estimate as primary.

Sensitivity check

The tenth eligible-provider question changed little

The four-provider primary panel has nine questions because Gemini P08 is missing. Claude and OpenAI completed all ten designed questions, allowing a valid all-available sensitivity for this eligible subset.

Nine-question primary and ten-question eligible-provider sensitivity
ProviderNine-question primaryTen-question sensitivityDifference
Claude10.99%10.95%-0.04 percentage points
OpenAI6.25%6.46%+0.21 percentage points

Including the tenth available question does not materially alter this descriptive Claude/OpenAI profile. It does not establish repeated-run stability: the study still contains one observed run per provider-question cell.

Hard measurement boundary

Why Gemini and Perplexity are excluded

Candidate-selection percentages require a comparable denominator. The locked observability contract treats the Claude and OpenAI exposed pools as complete for this analysis, while Gemini and Perplexity expose citation-biased source metadata.

Provider candidate-pool observability and selection eligibility
ProviderCandidate-pool statusCandidate-selection eligible?
ClaudeCompleteYes
GeminiCitation-biasedNo
OpenAICompleteYes
PerplexityCitation-biasedNo
Four-row observability matrix showing complete and eligible exposed candidate pools for Claude and OpenAI, and citation-biased ineligible metadata for Gemini and Perplexity.
Figure 2. Eligibility reflects observability of provider metadata, not model quality. A percentage for a citation-biased pool would use a different analytical denominator.
Open full-resolution figure
  • No four-provider rate

    Gemini and Perplexity percentages would look comparable while measuring a different denominator.

  • No hidden-state claim

    “Complete” is the status of the provider-exposed pool under the contract; it does not mean every internal retrieval or reasoning step is visible.

  • No universal breadth claim

    Candidate row totals reflect provider interfaces, search backends, orchestration and instrumentation.

Scale context

Raw reconciliation counts are not the headline estimand

Full frozen-snapshot reconciliation counts — not the question-macro estimand
ProviderExposed candidate rowsCited-source rowsCitation events
Claude66469118
OpenAI1,0426692
Separate Claude and OpenAI progressions from 664 and 1,042 exposed candidate rows to 69 and 66 cited-source rows, with citation events shown as a distinct object.
Figure 3. Full ten-question raw counts provide scale context only. The primary selection result is the equal-weight question-macro estimate in Figure 1, not 69 divided by 664 or 66 divided by 1,042.
Open full-resolution figure

Candidate rows, cited-source rows and citation events are distinct objects. The same cited source can support more than one citation event, and pooling rows would weight questions by candidate-pool size.

Interpretation

What candidate selection can—and cannot—tell us

  • Observed candidate membership

    For eligible stacks, the analysis records which sources appeared in the exposed pool associated with a response.

  • Final cited-source membership

    It identifies which observed candidates crossed into the visible cited-source set.

  • Within-question selection share

    It supports an equal-weight operational diagnostic across benchmark questions.

  • Panel sensitivity

    It shows whether adding the tenth completed Claude/OpenAI question materially changes the descriptive profile.

  • Not every internally considered source

    The study sees provider-exposed metadata, not hidden reasoning, training exposure or every retrieval stage.

  • Not the reason for selection

    The observed sequence does not identify a causal mechanism or isolate the effect of a page feature.

  • Not repeated-run probability

    One run per cell cannot estimate how often the same candidate would be selected again.

  • Not a provider ranking

    A lower fraction can arise from a larger exposed pool; a higher fraction can arise from a narrower one.

AI visibility diagnosis

A missing citation can reflect several different failures

  • A · Evidence does not exist publicly

    The relevant claim, proof point or page is absent from the public information environment.

  • B · Evidence is not observed in retrieval

    The page exists but does not appear in an eligible exposed candidate pool for the question.

  • C · Candidate is not selected

    The page appears in the pool, but another source enters the final cited set.

  • D · Citation maps to the wrong or weak claim

    The page is selected, but its role is not the one the company needs.

  • E · Citation does not support the claim well

    The source appears behind the answer but only partially supports—or fails to support—the statement.

The AI Answer Alignment funnel is exist → retrieve → select → support → represent. Reducing all five failure modes to “we were not cited” discards the diagnostic information needed to choose an intervention.

Company implications

Work on the stage where evidence is being lost

Candidate-stage visibility can distinguish a page that never entered the observable pool from one that entered but was not selected. That changes what a useful response looks like.

Exist
Create accurate, public evidence when the relevant product claim, proof or category explanation is missing.
Retrieve
Investigate question fit, accessibility and evidence discoverability when the page is absent from eligible exposed pools.
Select
Compare the retrieved page with selected alternatives for direct relevance, clarity and corroboration—without assuming causality.
Support and represent
Check that the citation maps to the material claim and contributes to an accurate, useful buyer answer.

Study design

The same fixed Cross-Model Citation Study

This companion is not a separate experiment or independent replication. The full benchmark contained 10 designed B2B buyer questions, 4 provider stacks, 40 expected provider-question cells, 39 successful responses and one missing Gemini response.

The primary candidate-selection comparison uses the nine-question complete-case panel for Claude and OpenAI. Because both eligible stacks completed all ten questions, it also reports a ten-question sensitivity. There is one observed run per provider-question cell.

Evidence boundaries

Limitations

  • Provider-interface constructs

    The pools are exposed metadata objects, not every internal document or reasoning step.

  • Observational sequence

    Candidate appearance followed by citation does not identify why selection occurred.

  • One run per cell

    The data cannot estimate repeated-run selection probability or model volatility.

  • Stack-specific construction

    Search backend, orchestration, instrumentation and answer-generation style can all shape exposed pools and rates.

  • Instructed protocol

    The candidate environments reflect the study's research workflow rather than unconstrained organic buyer usage.

Research record

Research and reproducibility

The public package contains provider observability rules, normalized candidate and source tables, candidate-selection outputs, benchmark-panel intervals, primary and all-available summaries, and reconciliation checks.

View the Cross-Model Citation Study publication package on GitHub.

FAQ

Frequently asked questions

Once a page appears in AI search results, how often is it cited?

For the two eligible provider stacks in this benchmark, the nine-question question-macro candidate-to-citation selection rate was approximately 11.0% for Claude and 6.2% for OpenAI.

Why are Gemini and Perplexity not included?

Their exposed source metadata is citation-biased rather than a validated complete candidate pool. A candidate-selection percentage would therefore use a different denominator and would not be comparable.

Does Claude select more sources than OpenAI?

Claude had a higher candidate-to-citation selection-rate point estimate in this benchmark. That does not mean Claude is better or universally more likely to cite a retrieved page. Candidate-pool construction and instrumentation differ.

How many candidate rows were exposed?

Across the full ten-question frozen snapshot, Claude had 664 candidate rows and OpenAI 1,042. These are exposed metadata-row counts, not universal retrieval-breadth estimates.

How many cited-source rows were there?

Across the full snapshot, Claude had 69 cited-source rows and OpenAI 66.

Why doesn't 69 divided by 664 equal the headline Claude rate exactly?

Because the headline rate is a question-macro average: selection is calculated within each question and then averaged equally across questions. The pooled row ratio gives greater weight to questions with larger candidate pools.

Did the tenth question change the result?

Very little. Claude was 10.99% in the nine-question primary panel and 10.95% in the ten-question sensitivity. OpenAI was 6.25% and 6.46%, respectively.

Does candidate selection tell us why a page was cited?

No. It tells us which observed candidates were selected, not the causal mechanism behind selection.

Is being retrieved enough for AI visibility?

Not if the goal is visible citation or accurate answer representation. Retrieval is one stage. A source may still need to be selected, attached to a relevant claim and used in a way that supports an accurate answer.

From benchmark to company evidence

See what this looks like for your company

Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.

See how Kojable works →
Explore AI citation monitoring →

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish