Kojable research · Cross-Provider AI Query Fan-Out Study · Companion analysis
What Query Fan-Out Can—and Cannot—Tell GEO Teams
A practical framework for using AI search-plan evidence without turning exploratory telemetry into false certainty
Query fan-out is useful because it exposes something traditional keyword research does not:
the family of retrieval needs an AI system may explore while answering one buyer question.
But fan-out telemetry is also easy to overinterpret.
A generated query can look precise enough to become a content brief. A shared query can look like proof of cross-model consensus. A provider-specific extension can look like a unique ranking opportunity. A late follow-up can look like evidence of deeper reasoning.
Our cross-provider study shows why those shortcuts are risky.
Claude, Gemini, OpenAI and Perplexity received the same ten B2B buyer questions and the same visible candidate-query scaffold. They differed materially in which candidates they selected, how much they moved beyond the scaffold, when extensions appeared, how much semantic overlap they shared, and which observable source pools they surfaced.
Those differences are informative.
They are not a complete theory of AI search.
Key conclusion
Query fan-out is best treated as observable process telemetry. It can help teams map recurring information needs, compare provider execution, identify shared and provider-specific search territory, and connect search plans to evidence surfaces.
It cannot, on its own, rank provider quality, reveal hidden reasoning, prove which content intervention will improve an AI answer, or predict the exact queries a provider will generate next time.
This final companion turns the study into a practical framework for GEO and AEO teams.
Read the flagship cross-provider study
- Query fan-out
- GEO measurement
- Evidence testing
- 11 min read
Measurement model
The five layers of AI answer discovery
A useful way to interpret fan-out research is to separate five layers:
- Buyer question
- Candidate or latent retrieval needs
- Observable search plan
- Observable evidence pool
- Final answer and business outcome
The study mainly observes layers 3 and 4, with controlled inputs at layers 1 and 2.
It does not directly identify hidden internal reasoning between those layers, and it does not experimentally manipulate content to measure layer-5 outcomes.
That distinction should shape how GEO teams use the results.

Supported descriptive uses
What query fan-out can tell you
1. Which information needs repeatedly appear around a buyer question
A buyer question is often broader than one search query.
Across the study, observable fan-out repeatedly touched needs such as:
- measurement and attribution;
- problem diagnosis;
- terminology expansion;
- solution or platform discovery;
- comparisons and alternatives;
- evidence and methodology;
- pricing and budget;
- implementation;
- technical fit;
- source and publisher verification.
That is useful for content strategy.
Instead of asking:
“What exact keyword should we rank for?”
a GEO team can ask:
“What information does an AI system appear to need in order to answer this buyer question well?”
That shift is important.
The practical unit becomes an information-need family, not one generated string.
2. How provider execution differs under the same scaffold
The four provider stacks did not use the same candidate space in the same way.
In the frozen experiment:
- OpenAI selected all supplied candidates;
- Perplexity selected a majority but frequently rewrote them;
- Claude selected fewer candidates and generated a substantial extension layer;
- Gemini selected the smallest success-conditional share and exposed the largest non-candidate-matched layer.
Those observations can help teams understand that:
“the fan-out queries” are not one universal list.
Provider-level execution can differ even when the buyer question and visible scaffold are held constant.
For a monitoring system, that means provider identity matters as a measurement dimension.
It does not mean every provider needs a separate content strategy.
3. Which needs are shared versus provider-specific
The semantic-clustering analysis found both:
- a cross-provider core;
- a substantial provider-specific tail.
That is a useful distinction.
A shared core can identify retrieval needs worth supporting broadly.
A provider-specific tail can identify areas to monitor for divergence.
But the consensus companion added an important qualification:
most cross-provider semantic consensus was seeded by the shared candidate space.
So a shared theme should be interpreted as:
“multiple providers used this semantic territory under the tested scaffold”
not:
“all providers independently discovered this as an unavoidable retrieval need.”
The practical implication is still useful, but the provenance matters.
4. Where query plans diverge from evidence pools
One of the clearest findings in the series is that query-plan similarity and source-pool similarity are different layers.
OpenAI and Perplexity sometimes had very similar semantic fan-out while surfacing almost entirely different URLs.
That means fan-out analysis can tell you:
what information needs are being explored
but not necessarily:
which exact source will satisfy them across every provider
For GEO teams, that supports a broader evidence strategy.
The relevant planning unit may be:
information need × provider × evidence surface
rather than one universal query → page mapping.
5. Where the observable sequence changes
Query order can reveal useful sequence structure.
The study found:
- Gemini often moved beyond the candidate-matched space early;
- Claude extensions were concentrated later;
- OpenAI’s small extension layer was entirely late;
- Claude showed directly observable result-conditioned follow-up cases.
This can help identify different search roles inside a plan:
- broad discovery;
- narrowing;
- verification;
- publisher targeting;
- author lookup;
- exact-page refinding.
That is more informative than counting queries alone.
Limits of inference
What query fan-out cannot tell you
1. Which provider is “best”
More queries are not automatically better.
More extensions are not automatically better.
A larger provider-specific tail is not automatically better.
A larger shared core is not automatically better.
A more specific query sequence is not automatically better.
Those are behavioural measures.
To rank provider quality, we would need downstream outcomes such as:
- answer accuracy;
- factual completeness;
- source authority;
- evidence relevance;
- citation support;
- cost;
- latency;
- reliability.
The present study does not combine those into a provider-quality benchmark.
2. What the provider “thought”
Observable query sequence is not chain-of-thought.
A late query does not prove deeper reasoning.
An early extension does not prove more creativity.
A publisher lookup does not reveal an internal decision rule.
The data show:
what was exported
not:
every internal step that generated it
This is especially important because provider telemetry differs.
Claude exposes direct single-query chronology. OpenAI and Perplexity expose batch-level structures. Gemini exposes a flatter run-level representation.
Those are not equivalent windows into hidden systems.
3. Whether an extension was independently invented
An extension is an operational provenance label.
It means the deterministic matcher could not connect the query to the supplied candidate space.
An extension might be:
- a genuinely new retrieval need;
- a narrow reformulation;
- an exact-title refinding query;
- a publisher check;
- a result-conditioned follow-up;
- a query that the deterministic matcher simply failed to map.
So:
extension ≠ creativity
and:
extension ≠ autonomous invention
That is why the series consistently uses language such as non-candidate-matched query.
4. What providers would do without the scaffold
This experiment was intentionally seeded.
That allowed us to compare:
- candidate selection;
- rewriting;
- expansion;
- consensus provenance.
But it also means the shared candidate space shaped the behaviour.
The consensus analysis made this especially clear: 97.6% of cross-provider clusters at the primary threshold were seeded consensus.
So this study does not estimate:
natural unseeded cross-provider fan-out
To answer that, we need a seeded-vs-unseeded experiment.
5. Which content change will improve AI visibility
This is the most important practical limitation.
The study did not manipulate:
- website content;
- third-party evidence;
- metadata;
- schema;
- source placement;
- citations;
- publisher coverage.
So it cannot establish:
“If you create content for this fan-out query, your AI visibility will improve.”
Nor can it establish:
“If you appear on this source, the provider will cite you.”
The study maps observable search behaviour.
It does not prove optimization effects.

Practical application
A practical GEO measurement framework
The research suggests a five-step workflow.
Step 1: Start with buyer questions, not keywords
Use real commercial questions such as:
- “How should I compare these vendors?”
- “Why is AI describing our company incorrectly?”
- “What proof should I trust?”
- “How should I measure AI visibility?”
- “When is a monitoring platform worth paying for?”
These are closer to the decisions AI systems are helping users make.
Step 2: Map information-need families
Cluster fan-out queries into recurring needs.
For example:
Buyer question: Which AI visibility platform is right for a B2B company?
Possible information needs:
- category terminology;
- vendor discovery;
- alternatives;
- pricing;
- measurement methodology;
- implementation;
- data/privacy fit;
- proof of effectiveness.
This prevents teams from creating a separate page for every generated string.
Step 3: Separate shared core from provider-specific tail
Ask:
- Which needs recur across providers?
- Which appear only in one provider trace?
- Which are candidate-derived?
- Which are non-candidate-matched?
- Which are stable across repeated runs?
The current experiment supports the first four.
Repeated runs are needed for the fifth.
Step 4: Connect needs to evidence surfaces
For each recurring information need, identify evidence that could plausibly satisfy it.
Examples:
| Information need | Possible evidence surface |
|---|---|
| Pricing | first-party pricing page, analyst comparison |
| Product capability | documentation, product page, credible review |
| Proof | case study, independent evaluation, customer evidence |
| Methodology | research report, technical methodology page |
| Security | trust center, documentation, third-party certification |
| Alternatives | category review, comparison page, industry publication |
These are strategy hypotheses, not experimentally proven prescriptions.
Step 5: Measure the final answer
The most important layer is the answer itself.
Did the AI system:
- describe the company accurately?
- mention the right category?
- include the relevant proof?
- avoid outdated claims?
- cite credible evidence?
- recommend the company in the right situations?
- distinguish it correctly from competitors?
Fan-out telemetry is most useful when connected to those answer outcomes.

Anti-patterns
The wrong way to use fan-out data
“We found 300 fan-out queries, so we need 300 pages”
No.
Many generated queries are:
- rewrites;
- near variants;
- verification searches;
- exact-page refinding;
- provider-specific tails.
The useful unit is often the underlying information need, not the literal string.
“Four providers searched this, so it must be important”
Not necessarily.
In a seeded experiment, shared candidates can manufacture apparent consensus.
Consensus should be decomposed by provenance.
“This provider generated more extensions, so it searches better”
No.
Extension count has no universal quality direction.
“This URL appeared in one provider, so we should get mentioned there”
Possibly, but that is a hypothesis.
Source appearance is not a demonstrated causal lever.
“The provider searched for a phrase, so that exact phrase should be optimized”
Not automatically.
The generated query may be one surface form of a broader retrieval need.
Evidence-aware use
The right way to use fan-out data
A stronger interpretation looks like this:
“Across several buyer questions, multiple provider traces repeatedly explored measurement methodology, vendor comparison and third-party proof. We should verify whether our evidence for those needs is complete, credible and discoverable, then retest final AI answers after improving the evidence.”
That is a defensible workflow.
It connects:
- observable search behaviour;
- evidence gaps;
- an intervention;
- a measurable answer outcome.
Research roadmap
How we would make this evidence stronger
The next generation of query-fan-out research should add:
- 30–50+ prompts
- several unrelated industries;
- seeded and unseeded conditions;
- randomized candidate order;
- 3–5 repeated runs;
- time-separated data collection;
- human validation of provenance;
- matched retrieval-depth analysis;
- direct answer-quality metrics;
- controlled content interventions.
The most important new outcome is not another query metric.
It is:
Does a measurable change in evidence coverage change the final AI answer?
That is the bridge from behavioural analysis to GEO optimization science.
Evidence hierarchy
A proposed GEO evidence ladder
A useful internal hierarchy is:
Level 1 — Observable
“What queries and sources appeared?”
Level 2 — Repeated
“Do they recur across runs and time?”
Level 3 — Cross-provider
“Do they recur across provider stacks?”
Level 4 — Intervention-linked
“Does changing relevant evidence alter retrieval or citation behaviour?”
Level 5 — Outcome-linked
“Does that change improve answer accuracy, positioning, recommendation or business impact?”
The current study is strongest at Levels 1–3.
It does not yet reach Levels 4–5.
That is where future GEO experiments should focus.

Conclusion
Conclusion
Query fan-out is valuable because it exposes the structure around a buyer question.
It can show:
- recurring information needs;
- provider-specific execution patterns;
- candidate adherence and extension;
- shared versus provider-specific semantic territory;
- observable search sequencing;
- source-pool divergence.
But it cannot, by itself, tell us:
- which provider is best;
- what a provider internally reasoned;
- which extensions were independently invented;
- what unseeded fan-out would look like;
- which content change will improve the final answer.
The practical value is in using fan-out as a measurement layer, not a universal optimization recipe.
A strong GEO workflow looks like:
buyer question → information need → observable fan-out → evidence coverage → final answer → intervention → retest
That keeps the strategy grounded in observable behaviour while preserving the distinction between:
what the system did
and
what we have actually proved will change the outcome.
FAQ
Frequently asked questions
Should companies optimize for every fan-out query?
No. Group related queries into information needs and evidence requirements before deciding whether new content is needed.
Is provider-specific fan-out important?
It can be useful for monitoring, but one provider-specific query in one run is not automatically a strategic priority.
Are shared fan-out queries more important?
They may identify useful common information needs, but in seeded designs the shared candidate scaffold can create much of the apparent consensus.
Does more fan-out mean better AI search?
No.
Can fan-out reveal what an AI model is thinking?
No. It reveals observable execution traces, not hidden chain-of-thought.
Can fan-out predict citations?
Not reliably from this study. Similar semantic plans often produced very different source pools.
Can fan-out tell me what content to create?
It can help identify information gaps and hypotheses. Controlled interventions and retesting are needed to establish whether a content change improves AI answers.
What is the most useful unit for GEO planning?
Usually the recurring buyer information need, connected to credible evidence and tested against final AI answers.
What should teams measure after fan-out?
Answer accuracy, positioning, recommendation context, evidence quality, citation support and change after intervention.
