Kojable research · Cross-Model Citation Study
From Search Results to AI Citations: What Claude and OpenAI Selected
Appearing in an AI system's exposed search or candidate pool is not the same thing as being cited in the final answer.
Key finding
Across the nine-question matched panel, the question-macro candidate-to-citation selection rate was approximately 11.0% for Claude and 6.2% for OpenAI. Most observed candidates did not become final citations.
Qualification: Candidate-selection analysis is valid only for Claude and OpenAI under the locked observability contract. Gemini and Perplexity expose citation-biased metadata rather than comparable complete candidate pools. These rates are descriptive operational diagnostics—not provider-quality scores, causal probabilities or repeated-run estimates.
- 2 eligible provider stacks
- 9-question primary panel
- 10-question sensitivity
- 10 min read
Answer first
How often did exposed candidates become citations?
Direct answer
Once a source appeared in an eligible exposed candidate pool, only a minority of those observed candidates were ultimately selected for citation: approximately 11.0% for Claude and 6.2% for OpenAI in the nine-question question-macro comparison.
These are equal-weight question-macro estimates, not pooled raw-row ratios. They do not rank provider quality, and comparable rates are not reported for Gemini or Perplexity because their exposed source denominators are citation-biased.
| Provider | Nine-question primary | Benchmark-panel interval | Ten-question sensitivity |
|---|---|---|---|
| Claude | 10.99% | 7.7%–14.5% | 10.95% |
| OpenAI | 6.25% | 5.3%–7.5% | 6.46% |
The locked estimates are 10.99208219% and 6.24642793% in the primary panel, with all-available sensitivities of 10.94550555% and 6.45511847%, respectively. Display values are rounded for readability.
Stage model
Retrieval exposure and citation are different states
The exposed metadata supports a stage model without claiming access to hidden model state.
- Not observed in the exposed candidate pool
The source is absent from the provider-exposed search or candidate records associated with the response.
- Observed as a candidate
The page enters the observable retrieval environment, but this does not show that it influenced the answer.
- Selected for citation
The candidate appears in the final cited-source set and becomes visible answer evidence.
- Mapped to a claim
A citation event can be connected to a claim or answer segment.
- Evaluated for support
A separate evidence layer asks whether the cited source actually supports the linked claim.
The concise sequence is candidate exposure → citation selection → claim mapping → evidentiary support. This companion focuses on the observable transition from candidate exposure to final citation.
Finding 1
Claude and OpenAI selected a minority of observed candidates

Claude's point estimate was higher in this observed panel, but the central finding is not a provider winner. Candidate exposure and citation selection were materially different stages in both eligible stacks.
Estimand
Why the primary rate is question-macro
The headline calculation first measures selection inside each eligible provider-question cell and then averages questions with equal weight. This prevents a question with a large exposed pool from automatically dominating the provider summary.
- Question-macro result
- Answers: for a typical question in this fixed benchmark, what share of its exposed candidate pool became cited sources?
- Pooled row ratio
- Answers: what fraction of all candidate rows across all questions became cited rows? Questions with larger pools receive more weight.
- Primary analytical unit
- The question. Provider-question responses, candidate sources, cited sources and citation events are nested objects.
For example, selection rates of 20% from 20 candidates and 5% from 200 candidates produce a 12.5% equal-weight macro average but a 6.4% pooled ratio. Neither calculation is inherently wrong; they answer different questions. This study specifies the question-macro estimate as primary.
Sensitivity check
The tenth eligible-provider question changed little
The four-provider primary panel has nine questions because Gemini P08 is missing. Claude and OpenAI completed all ten designed questions, allowing a valid all-available sensitivity for this eligible subset.
| Provider | Nine-question primary | Ten-question sensitivity | Difference |
|---|---|---|---|
| Claude | 10.99% | 10.95% | -0.04 percentage points |
| OpenAI | 6.25% | 6.46% | +0.21 percentage points |
Including the tenth available question does not materially alter this descriptive Claude/OpenAI profile. It does not establish repeated-run stability: the study still contains one observed run per provider-question cell.
Hard measurement boundary
Why Gemini and Perplexity are excluded
Candidate-selection percentages require a comparable denominator. The locked observability contract treats the Claude and OpenAI exposed pools as complete for this analysis, while Gemini and Perplexity expose citation-biased source metadata.
| Provider | Candidate-pool status | Candidate-selection eligible? |
|---|---|---|
| Claude | Complete | Yes |
| Gemini | Citation-biased | No |
| OpenAI | Complete | Yes |
| Perplexity | Citation-biased | No |

- No four-provider rate
Gemini and Perplexity percentages would look comparable while measuring a different denominator.
- No hidden-state claim
“Complete” is the status of the provider-exposed pool under the contract; it does not mean every internal retrieval or reasoning step is visible.
- No universal breadth claim
Candidate row totals reflect provider interfaces, search backends, orchestration and instrumentation.
Scale context
Raw reconciliation counts are not the headline estimand
| Provider | Exposed candidate rows | Cited-source rows | Citation events |
|---|---|---|---|
| Claude | 664 | 69 | 118 |
| OpenAI | 1,042 | 66 | 92 |

Candidate rows, cited-source rows and citation events are distinct objects. The same cited source can support more than one citation event, and pooling rows would weight questions by candidate-pool size.
Interpretation
What candidate selection can—and cannot—tell us
- Observed candidate membership
For eligible stacks, the analysis records which sources appeared in the exposed pool associated with a response.
- Final cited-source membership
It identifies which observed candidates crossed into the visible cited-source set.
- Within-question selection share
It supports an equal-weight operational diagnostic across benchmark questions.
- Panel sensitivity
It shows whether adding the tenth completed Claude/OpenAI question materially changes the descriptive profile.
- Not every internally considered source
The study sees provider-exposed metadata, not hidden reasoning, training exposure or every retrieval stage.
- Not the reason for selection
The observed sequence does not identify a causal mechanism or isolate the effect of a page feature.
- Not repeated-run probability
One run per cell cannot estimate how often the same candidate would be selected again.
- Not a provider ranking
A lower fraction can arise from a larger exposed pool; a higher fraction can arise from a narrower one.
AI visibility diagnosis
A missing citation can reflect several different failures
- A · Evidence does not exist publicly
The relevant claim, proof point or page is absent from the public information environment.
- B · Evidence is not observed in retrieval
The page exists but does not appear in an eligible exposed candidate pool for the question.
- C · Candidate is not selected
The page appears in the pool, but another source enters the final cited set.
- D · Citation maps to the wrong or weak claim
The page is selected, but its role is not the one the company needs.
- E · Citation does not support the claim well
The source appears behind the answer but only partially supports—or fails to support—the statement.
The AI Answer Alignment funnel is exist → retrieve → select → support → represent. Reducing all five failure modes to “we were not cited” discards the diagnostic information needed to choose an intervention.
Company implications
Work on the stage where evidence is being lost
Candidate-stage visibility can distinguish a page that never entered the observable pool from one that entered but was not selected. That changes what a useful response looks like.
- Exist
- Create accurate, public evidence when the relevant product claim, proof or category explanation is missing.
- Retrieve
- Investigate question fit, accessibility and evidence discoverability when the page is absent from eligible exposed pools.
- Select
- Compare the retrieved page with selected alternatives for direct relevance, clarity and corroboration—without assuming causality.
- Support and represent
- Check that the citation maps to the material claim and contributes to an accurate, useful buyer answer.
Study design
The same fixed Cross-Model Citation Study
This companion is not a separate experiment or independent replication. The full benchmark contained 10 designed B2B buyer questions, 4 provider stacks, 40 expected provider-question cells, 39 successful responses and one missing Gemini response.
The primary candidate-selection comparison uses the nine-question complete-case panel for Claude and OpenAI. Because both eligible stacks completed all ten questions, it also reports a ten-question sensitivity. There is one observed run per provider-question cell.
Evidence boundaries
Limitations
- Provider-interface constructs
The pools are exposed metadata objects, not every internal document or reasoning step.
- Observational sequence
Candidate appearance followed by citation does not identify why selection occurred.
- One run per cell
The data cannot estimate repeated-run selection probability or model volatility.
- Stack-specific construction
Search backend, orchestration, instrumentation and answer-generation style can all shape exposed pools and rates.
- Instructed protocol
The candidate environments reflect the study's research workflow rather than unconstrained organic buyer usage.
Research record
Research and reproducibility
The public package contains provider observability rules, normalized candidate and source tables, candidate-selection outputs, benchmark-panel intervals, primary and all-available summaries, and reconciliation checks.
View the Cross-Model Citation Study publication package on GitHub.
FAQ
Frequently asked questions
Once a page appears in AI search results, how often is it cited?
For the two eligible provider stacks in this benchmark, the nine-question question-macro candidate-to-citation selection rate was approximately 11.0% for Claude and 6.2% for OpenAI.
Why are Gemini and Perplexity not included?
Their exposed source metadata is citation-biased rather than a validated complete candidate pool. A candidate-selection percentage would therefore use a different denominator and would not be comparable.
Does Claude select more sources than OpenAI?
Claude had a higher candidate-to-citation selection-rate point estimate in this benchmark. That does not mean Claude is better or universally more likely to cite a retrieved page. Candidate-pool construction and instrumentation differ.
How many candidate rows were exposed?
Across the full ten-question frozen snapshot, Claude had 664 candidate rows and OpenAI 1,042. These are exposed metadata-row counts, not universal retrieval-breadth estimates.
How many cited-source rows were there?
Across the full snapshot, Claude had 69 cited-source rows and OpenAI 66.
Why doesn't 69 divided by 664 equal the headline Claude rate exactly?
Because the headline rate is a question-macro average: selection is calculated within each question and then averaged equally across questions. The pooled row ratio gives greater weight to questions with larger candidate pools.
Did the tenth question change the result?
Very little. Claude was 10.99% in the nine-question primary panel and 10.95% in the ten-question sensitivity. OpenAI was 6.25% and 6.46%, respectively.
Does candidate selection tell us why a page was cited?
No. It tells us which observed candidates were selected, not the causal mechanism behind selection.
Is being retrieved enough for AI visibility?
Not if the goal is visible citation or accurate answer representation. Retrieval is one stage. A source may still need to be selected, attached to a relevant claim and used in a way that supports an accurate answer.
From benchmark to company evidence
See what this looks like for your company
Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.
