Kojable research · Cross-Provider AI Query Fan-Out Study
How Claude, Gemini, OpenAI and Perplexity Fan Out the Same Buyer Questions
The same buyer question does not necessarily become the same observable web-search plan inside different AI provider stacks.
Key finding
The four provider stacks did not expose one common query plan. They differed in candidate selection, non-candidate-matched query generation, context novelty, query sequencing and core-versus-provider-specific composition.
Qualification: This fixed exploratory benchmark covers ten designed questions, a shared candidate-query scaffold and one observed run per provider-question cell. The results are behavioural diagnostics, not provider-quality rankings.
- 4 provider stacks
- 10 buyer questions
- Shared candidate scaffold
- 18 min read
Executive summary
One input protocol produced several observable search-plan patterns
Kojable gave Claude, Gemini, OpenAI and Perplexity the same ten designed B2B buyer questions, the same task context and the same visible space of 100 candidate queries. The analysis then compared which supplied candidates appeared, which queries moved beyond that space, how sequences developed, where providers converged semantically and how much their observable source pools overlapped.
Across 40 intended provider-question cells, 39 runs produced query-bearing exports; Gemini P08 failed and was not imputed. The resulting dataset contains 361 observable query strings and 203 prompt-local semantic clusters. Queries, search events, results, URLs and citations are nested observations. The buyer question—not the individual query row—is the matched inferential unit.
Candidate use and non-candidate-matched query shares differed sharply, but neither direction is inherently better. Most cross-provider semantic agreement was connected to the shared candidate scaffold, and the comparable provider stacks still surfaced largely different URLs and domains. Query similarity and evidence similarity are therefore separate measurement layers.
Answer first
Do major AI providers fan out the same buyer question in the same way?
Direct answer
Given the same buyer question, task instructions and supplied candidate-query space, Claude, Gemini, OpenAI and Perplexity exposed materially different candidate-selection, query-extension, sequencing and semantic-plan patterns. Most cross-provider semantic consensus was seeded by the shared candidate scaffold rather than emerging independently.
These differences are exploratory behavioural observations, not evidence that one provider searched better.
- Candidate selection
Providers selected different shares of the common supplied query space.
- Extension and vocabulary
Non-candidate-matched query generation and context-novel vocabulary differed substantially.
- Plan composition
Shared semantic cores coexisted with provider-specific tails of very different sizes.
- Evidence surfaces
Similar query directions did not guarantee similar retrieved or cited source pools.
Study design
A fixed, seeded and matched cross-provider benchmark
- Provider stacks
- Claude, Gemini, OpenAI and Perplexity.
- Prompt design
- Ten designed B2B buyer questions across diagnosis, commercial evaluation, solution evaluation, measurement, governance and implementation intents.
- Cells and runs
- 40 intended provider-question cells; 39 successful query-bearing runs; Gemini P08 failed; one observed run per cell.
- Seed space
- 100 supplied candidate queries shared across providers. Candidate order was not randomized.
- Observed inventory
- 361 query strings grouped into 203 prompt-local semantic clusters.
- Inferential unit
- The buyer question. Provider-question runs, queries, search events or batches, results, URLs, domains and citations are nested observations.
Provider observability is not uniform
| Provider | Observable query structure | Result linkage | Main comparability constraint |
|---|---|---|---|
| Gemini | Flat ordered query list | Run-level | No native batch identity; exported destinations remain unresolved redirect wrappers |
| OpenAI | Ordered multi-query batches | Batch-level | Individual result attribution within a batch is unavailable |
| Claude | Ordered single-query tool calls | Direct query-level | Clearest observable chronological follow-up evidence |
| Perplexity | Ordered multi-query batches | Batch-level | Individual result attribution within a batch is unavailable |
An observable query count is not a universal backend search-request count. The study therefore does not rank provider search effort, search depth or retrieval efficiency from raw query totals.
Finding 1
The providers used the same candidate space differently
| Provider | Selected candidates | Successful-run candidate opportunities | Success-conditional selection rate |
|---|---|---|---|
| Gemini | 32 | 90 | 35.6% |
| Claude | 49 | 100 | 49.0% |
| Perplexity | 67 | 100 | 67.0% |
| OpenAI | 100 | 100 | 100.0% |
The frozen aggregate table records Gemini as 32 selected candidates out of 100 because that denominator includes ten opportunities attached to the failed P08 response. The public comparison conditions on the nine successful Gemini runs: 32 of 90, or 35.6%. A failed response is missing data, not ten observed non-selections.
Candidate-selection rate has no universal quality direction. More candidate use may reflect breadth or weak filtering; less may reflect prioritisation or missed retrieval paths. The experiment measures observable selection behaviour.
Finding 2
Non-candidate-matched query generation differed even more
An extension is shorthand for an observed query that cannot be connected to the supplied candidate space under the deterministic candidate and provenance rules. It is an operational label, not a claim of creativity, intelligence, autonomy, originality, hidden reasoning or retrieval quality.
| Provider | Observable queries | Non-candidate-matched queries | Extension share | Context-novel token share |
|---|---|---|---|---|
| Gemini | 88 | 54 | 61.4% | 17.9% |
| Claude | 84 | 36 | 42.9% | 12.6% |
| OpenAI | 118 | 14 | 11.9% | 0.2% |
| Perplexity | 71 | 3 | 4.2% | 0.0% |
Context-novel vocabulary means meaningful query tokens absent from the visible buyer prompt, task instructions, candidate queries and earlier observable queries. It is not “new knowledge”; hidden provider context remains possible.
- Different need or decomposition
A non-candidate-matched query may introduce a genuinely different need, specialise an earlier need or split it into smaller searches.
- Reformulation or refinding
It may restate an earlier need, refind an exact title or page, or target a publisher or author.
- Operational mismatch
It may reflect hidden provider context or fall outside the deterministic matching rules.
Finding 3
Most cross-provider consensus was already seeded
The analysis created 203 prompt-local semantic clusters before assigning candidate provenance. That order matters: candidate origin did not boost clustering similarity.
| Consensus class | Cluster count | Interpretation |
|---|---|---|
| Seeded consensus | 82 | Cross-provider agreement connected to the supplied candidate space |
| Emergent consensus | 1 | Extension-only cross-provider consensus; a case study, not a general rate |
| Mixed consensus | 1 | Cluster containing both candidate-derived and extension queries |
| Provider-specific | 119 | Semantic cluster observed in only one provider stack |
Most observed cross-provider semantic agreement was connected to the shared candidate-query scaffold. The results do not show that four independent AI systems naturally discovered the same query plan; emergent extension-only consensus was extremely rare in this seeded benchmark.
Finding 4
A shared core sat beside different provider-specific tails
| Provider | Core query share | Provider-specific query share |
|---|---|---|
| Gemini | 29.5% | 64.8% |
| OpenAI | 40.7% | 28.8% |
| Claude | 44.0% | 46.4% |
| Perplexity | 62.0% | 4.2% |
These shares describe convergence and divergence in observable search plans. A larger core is not superior, a larger tail is not innovative, a smaller tail is not efficient, and provider-specific share is not a provider-quality score.
Recurring query families included measurement and attribution, solution discovery, terminology expansion, problem diagnosis, comparison, evidence and methodology, implementation, commercial evaluation and source targeting. Gemini frequently exposed exact-page or quoted-concept refinding forms, while Claude frequently exposed publisher, creator and author targeting. These are cautious examples, not a provider-personality taxonomy; no post-hoc significance sweep was run across all query families.
Finding 5
Query order revealed different observable timing patterns
Observable sequence means exported execution chronology, not a record of hidden reasoning depth. Counts remain essential because extension volume differed greatly between providers.
| Provider | Extension queries | Early | Middle | Late |
|---|---|---|---|---|
| Claude | 36 | 0.0% | 30.6% | 69.4% |
| Gemini | 54 | 25.9% | 31.5% | 42.6% |
| OpenAI | 14 | 0.0% | 0.0% | 100.0% |
| Perplexity | 3 | 33.3% | 0.0% | 66.7% |
Extensions appeared in all nine successful Gemini runs, all ten Claude runs, seven OpenAI runs and three Perplexity runs. The prompt-paired exploratory analysis supported an earlier first-extension position for Gemini than Claude. Perplexity has only three extension queries and cannot support a strong timing interpretation.
Observable chronology
Eight Claude queries show result-conditioned observable follow-up
Claude supplied the clearest query-to-result chronology in the dataset. Eight directly observable extension queries across four prompts contain evidence connecting an earlier retrieved result, distinctive vocabulary, a later extension query and its subsequent result set.
- P01 · 2 cases
- Later searches targeted Search Engine Land and Sword and the Script vocabulary introduced by prior results.
- P05 · 4 cases
- Later searches used an exact title phrase, the name Alex Birkett, SparkToro title wording and a Search Engine Land title/byline refind.
- P07 · 1 case
- A later query used the distinctive Forrester title phrase “Seven Roles, One Goal”.
- P08 · 1 case
- A later query used Kevin Indig and Growth Memo vocabulary first visible in a prior result.
This supports result-conditioned observable follow-up in those Claude traces. It does not prove that every Claude extension is adaptive, that Claude generally reasons more deeply, that the same process occurs in all Claude runs, that other providers do not adapt, or that hidden chain-of-thought has been observed.
Finding 6
Similar query plans did not imply similar source pools
Gemini is excluded from URL and domain overlap interpretation because its exported destinations remain unresolved provider redirect wrappers. Its zero rows in the raw summary are not evidence of no shared web destinations.
| Provider pair | Retrieved URL Jaccard | Retrieved domain Jaccard | Cited URL Jaccard | Cited domain Jaccard |
|---|---|---|---|---|
| Claude–Perplexity | 0.037 | 0.084 | 0.009 | 0.017 |
| OpenAI–Claude | 0.014 | 0.045 | 0.018 | 0.043 |
| OpenAI–Perplexity | 0.006 | 0.040 | 0.012 | 0.018 |
Query overlap and evidence overlap are different layers
- Candidate selection and semantic plans
Candidate-selection Jaccard and semantic-cluster Jaccard were positively associated: Spearman ρ=0.682, with a prompt-cluster bootstrap interval of 0.354 to 0.908. Providers selecting more of the same supplied candidates tended to expose more similar semantic plans in this sample; the relationship is observational and scaffold-dependent.
- Retrieved URLs and domains
Semantic-cluster overlap versus retrieved URL overlap was inconclusive (ρ=-0.108; interval -0.498 to 0.251), as was retrieved domain overlap (ρ=-0.184; interval -0.515 to 0.137).
- Cited URLs
The cited-URL association was positive and exploratory (ρ=0.288; interval 0.049 to 0.547), not a stable general relationship.
- Cited domains
The cited-domain association remained inconclusive (ρ=0.215; interval -0.098 to 0.538).
Negative evidence
Several plausible differences remained inconclusive
- Prompt-facet coverage
The study did not resolve robust adjusted provider-pair differences.
- Search-stage and query-family switching
Neither switching analysis resolved robust adjusted provider-pair differences.
- Near-repeat rate
The stronger exploratory criteria were not met.
- Native-batch characteristics
OpenAI and Perplexity did not show separately adjusted differences across the four analysed characteristics.
- Perplexity within-provider comparison
Only three prompts contained both candidate-only and extension-containing batches, below the formal-test threshold.
Inconclusive does not mean equivalent. The sample and observed zero-heavy metrics may simply be unable to distinguish small or unstable effects.
AEO and GEO implications
Plan for retrieval needs and evidence surfaces, not literal query strings
- Plan around information needs
Test whether the company has clear evidence for recurring families such as measurement, diagnosis, comparison, methodology, implementation and commercial evaluation—not whether it has one page per exported query.
- Separate the common core from provider-specific tails
Track recurring cross-provider needs, provider-specific paths and how those paths change over time.
- Treat source visibility as a multi-surface problem
Evidence visible in one provider environment may not transfer to another. Test first-party proof, technical documentation and credible third-party surfaces across stacks.
- Do not build a provider leaderboard
More queries, more extensions, a larger core or a larger provider-specific tail are behavioural telemetry until linked to downstream answer and evidence quality.
Within Kojable's AI Answer Alignment framing, the observable diagnostic path is exist / evidence → retrieval need → observable query path → source pool → answer representation. It is a measurement model, not a deterministic causal chain.
The study does not support creating one page per fan-out query, optimising for every literal exported query, copying one provider's plan or assuming that source visibility transfers between providers.
Exploratory inference
The statistical layer was created after descriptive analysis
All inferential results are exploratory and hypothesis-generating, not confirmatory. The scientific value lies in broad patterns recurring across matched prompts, not in treating every qualifying contrast as an independent discovery.
- Friedman omnibus tests
Applied across complete prompt blocks.
- Exact paired sign-flip tests
Used for provider-pair comparisons at the prompt level.
- Prompt-paired bootstrap intervals
Preserved the buyer question as the resampling unit.
- Benjamini–Hochberg adjustment
Controlled the false-discovery rate across related tests.
- Leave-one-prompt-out checks
Tested whether effect directions depended on a single prompt.
Rule for a stronger exploratory signal
- Interval
The effect interval excluded zero.
- Adjusted threshold
Benjamini–Hochberg adjusted q was below 0.05.
- Directional stability
Leave-one-prompt-out direction remained stable.
Twenty-six provider-pair contrasts met that rule. They are not “26 discoveries”; they largely represent related dimensions of candidate use, query expansion, novelty, core-versus-tail composition, plan breadth and sequence characteristics.
Evidence boundaries
Limitations
- Ten designed prompts
They are not a random sample of all buyer questions.
- One run per provider-question cell
The study cannot estimate run-to-run stochastic variability.
- Shared candidate seeding
Most consensus is scaffold-dependent; an unseeded control is required.
- Candidate order not randomized
Position may affect selection and is entangled with candidate quality and ordering structure.
- Heterogeneous telemetry
Run-, batch- and query-level result structures are not equivalent across providers.
- Operational extension label
Non-candidate-matched does not mean an invented information need.
- Taxonomy and clustering error
The deterministic, auditable rules still contain modelling choices.
- Unresolved Gemini URL wrappers
Gemini is excluded from cross-provider URL-overlap interpretation.
- Post-descriptive statistical plan
Inference is exploratory, not confirmatory.
- No downstream quality outcome
The study does not rank accuracy, usefulness, source authority, cost, latency, reliability or business impact.
Research agenda
What should be tested next
The next design should estimate whether provider behaviour changes differently when the same candidate scaffold is present versus absent.
- Broader prompt panel
Use 30–50 or more prompts across unrelated domains.
- Seeded and unseeded conditions
Separate natural convergence from scaffold-induced agreement.
- Randomized candidate order
Estimate position effects independently from candidate content.
- Repeated, time-separated runs
Collect three to five runs per provider-question-condition cell in multiple time blocks.
- Stronger chronology
Capture direct query-result linkage where technically available.
- Resolved destinations
Resolve redirect wrappers before cross-provider source comparison.
- Human validation
Review candidate matching and semantic clustering.
- Downstream outcomes
Measure answer and evidence quality, cost, latency and reliability.
Research record
Research and reproducibility
The publication package preserves the frozen data, normalised events, provenance rules, query taxonomy, consensus outputs, sequencing outputs, overlap outputs, statistical tables, finding registry, reports and figures used to support the public claims.
The technical report, evidence appendix and canonical finding registry are retained as part of that package.
Public repository access is currently unavailable. Until the public repository route is restored, this page does not claim that readers can independently access the complete reproducibility package online.
FAQ
Frequently asked questions
What is a fan-out query?
In this study, a fan-out query is an observable web-search query exposed during an AI provider's attempt to answer a buyer question. A single buyer question can lead to several related queries covering different information needs.
Which provider generated the most non-candidate-matched queries?
Gemini had the highest observed extension share in this benchmark at 61.4% of its observable queries, followed by Claude at 42.9%, OpenAI at 11.9% and Perplexity at 4.2%. This is an operational provenance measure, not a quality ranking.
Did the four providers independently discover the same queries?
Mostly not in this seeded experiment. Of 203 semantic clusters, 82 were seeded-consensus clusters, while only one was classified as emergent consensus and one as mixed consensus. Most cross-provider semantic agreement was therefore connected to the common candidate space.
Does a larger provider-specific tail mean better search?
No. A provider-specific query can introduce useful evidence, but it can also represent reformulation, refinding, redundant search or a different exporter style. The experiment does not assign a universal quality direction to tail size.
Did similar query plans retrieve the same websites?
Not usually in the comparable source pools. Across OpenAI, Claude and Perplexity, mean retrieved-URL Jaccard overlap was only about 0.006 to 0.037. Semantic-query overlap was not strongly associated with retrieved URL or domain overlap in this small sample.
Does this prove that companies should create a page for every fan-out query?
No. The evidence supports testing broader information-need coverage, not manufacturing a separate page for every observed query string. Many queries are semantically related, provider-specific or influenced by the supplied candidate scaffold.
Which provider was best?
This study does not answer that question. It measures observable search-plan behaviour, not overall answer quality, accuracy, source authority, cost, latency or business impact.
Are these stable provider behaviours?
Not yet established. Each provider-question cell contains only one observed run. Repeated, time-separated runs are required before these patterns can be treated as stable provider traits.
From benchmark to company evidence
See what this looks like for your company
Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.