Kojable research · Narrative fidelity · Article 2

Citation Rank and Semantic Alignment: Why Shared Sources Don’t Make AI Answers Converge

Published Updated By Piush Vaish

Article 1 established the broad pattern. This technical continuation explains why shared Rank-1 sources did not make whole finance answers meaningfully more similar, even though Rank 1 received a small within-response semantic-alignment premium.

Key finding

Rank 1 showed a modest local alignment association without producing global answer convergence.

Qualification: This observational analysis does not establish that changing citation order would cause semantic alignment to change.

  • Narrative fidelity
  • Semantic alignment
  • Citation rank
  • Article 2 of 5
  • 13 min read
Two-panel comparison showing nearly equal whole-answer similarity for shared Rank-1 and lower-ranked sources, alongside a modest Rank-1 alignment premium over shuffled ranks.
Figure 1. Shared-source whole-answer similarity and within-response semantic alignment operate at different levels. The similarity difference was effectively null, while Rank 1 received a modest alignment premium among eligible answers.
Open full-resolution figure
0.0035shared-source similarity differenceUnsupportedThe 95% interval crossed zero.
+1.47ppRank-1 alignment versus shuffled expectation
660eligible alignment responses
0.484directional p-value for shared-source similarityNull-compatible

Technical continuation

Article 1 established the pattern; Article 2 explains the tension

Article 1 of the Narrative Fidelity series separated four possible citation-rank outcomes across 1,500 generated finance responses. It found a modest semantic-alignment advantage and a clearer entity-visibility association at Rank 1, but no credible whole-answer similarity or recommendation effect.

This page addresses the apparent tension inside that result. How can Rank 1 align slightly more strongly with an answer without making answers that share the same Rank-1 source converge?

The answer is that the analyses estimate different relationships: whole-answer convergence versus within-answer evidence reflection.

Answer first

Can Rank 1 align locally without controlling the whole answer?

  • A modest local signal, not narrative control

    Rank 1 was associated with a small within-response semantic-alignment premium. Sharing a Rank-1 source did not make complete answers meaningfully more similar. Those findings are compatible because one concerns local evidence reflection and the other concerns global answer convergence.

Rank 1 is associated with a small within-response semantic-alignment premium, but sharing a Rank-1 source does not make whole answers meaningfully more similar.

Estimands

Two questions that sound similar but are not

The distinction begins with the unit of analysis.

  • Global convergence compares answers with answers

    If two responses cite the same source at Rank 1, are those two complete responses more semantically similar than a lower-ranked shared-source comparison group?

  • Local reflection compares evidence within one answer

    Given the evidence opportunities inside a response, does the source at Rank 1 receive more alignment share than it would under random reassignment of those same source ranks?

Treating these estimands as interchangeable would turn a modest local association into an unsupported claim about global narrative control.

Finding 1 · Unsupported effect

Shared Rank-1 sources did not make whole answers more similar

The Rank-1 comparison included 155 response pairs across 64 source-and-base-query cells. The lower-rank comparison included 200 pair occurrences—187 unique pairs—across 106 source-and-query cells.

Source/query cells received equal weight so that a highly common source could not dominate the estimate simply by producing many overlapping response pairs.

Shared Rank-1 source
Mean whole-answer cosine similarity: 0.7247.
Shared Rank 3-or-lower source
Mean whole-answer cosine similarity: 0.7213.
Estimated difference
0.0035; 95% interval −0.0199 to 0.0253.
Inference
Directional permutation p-value 0.484; standardized effect approximately 0.045.

This is an effectively null result. The interval remains compatible with a small negative effect, no effect or a small positive effect. It does not support the claim that sharing the same first cited source meaningfully increases whole-answer similarity.

Interpretation

Why whole-answer similarity can stay flat

A shared source is only one input to a generated response. Two answers can cite the same source first and still diverge because of:

  • Other evidence

    The answer can include different sources alongside the shared source.

  • Claim selection

    Each response can emphasize different claims from the same source.

  • Answer structure

    Organization and level of detail can remain substantially different.

  • User need

    Scope, intent and domain context can change what the answer prioritizes.

  • Generative variability

    Framing, synthesis and wording can vary across generations.

The null similarity result does not mean that the source contributes nothing. Whole-answer similarity is a demanding outcome: it asks whether two complete responses converge, not whether a particular piece of evidence leaves a detectable footprint inside each answer.

Finding 2 · Modest association

Rank 1 showed a modest within-response alignment advantage

The second analysis treated each response as its own evidence opportunity set. Evidence text was divided into chunks, response text into sections, and the strongest chunk-to-section similarities were summarized before sources were combined into the rank bins present in that response.

Exact source ranks were then shuffled within the same response while its answer, source set and source-level semantic scores stayed fixed.

Among 660 eligible responses, observed Rank-1 alignment share was 0.4043. The shuffled expectation was 0.3896. The difference was 0.0147, equivalent to 1.47 percentage points.

The 95% response-block interval was 0.0052 to 0.0241. The directional permutation p-value was 0.001, and the two-sided p-value was 0.004.

The association is statistically discernible but modest. It does not show that Rank 1 dominates the narrative, and it does not establish that moving a source into first cited position would cause stronger alignment.

Eligibility

The opportunity set is part of the result

The 1.47-point estimate only has meaning where Rank 1 had usable lower-ranked evidence to compete against. The primary alignment population required:

  • Competing sources

    At least two usable cited sources had to be present.

  • Rank-1 evidence

    The first cited-source position needed usable evidence.

  • Lower-rank opportunity

    At least one usable lower-rank bin had to be available.

  • Comparable text

    The evidence needed enough chunks and the answer enough sections to calculate alignment.

The main outcomes were 660 analyzable responses, 343 with fewer than two usable sources, 89 without usable Rank-1 evidence and 5 with fewer than two usable rank bins.

Coverage diagram showing 1,500 total responses and the narrower alignment opportunity set: 660 analyzable, 343 with too few usable sources, 89 without usable Rank-1 evidence and 5 with too few rank bins.
Figure 2. The alignment estimate applies only where an answer contains a genuine comparison between usable Rank-1 and lower-ranked evidence. The eligibility restriction defines the population to which the estimate applies.
Open full-resolution figure

Scale

What a 1.47-point alignment premium actually means

The observed Rank-1 share was approximately 40.4%, while the response-specific shuffled expectation was already about 39.0%.

It would be incorrect to compare 40.4% with an artificial 25% four-bin baseline. Not every response contained four bins, and the number of usable sources per bin varied. The valid counterfactual preserves each response’s actual evidence opportunities and randomly reassigns the observed ranks.

The finding is not “Rank 1 gets 40% of every answer.” It is that, among eligible answers, Rank 1 received a slightly greater share of semantic alignment than expected under within-response rank shuffling. The effect is best understood as a modest tilt.

Central distinction

Local evidence reflection is not global narrative convergence

Two answers can share the same first cited source while incorporating different additional evidence, choosing different claims, organizing those claims differently, serving different user needs and varying in framing or wording.

The shared source can therefore leave a modest local footprint without imposing a shared whole-answer narrative.

Interpretation

Narrative fidelity can increase locally without producing answer-level convergence.

Citation position is not meaningless, but it is not a reliable proxy for narrative control.

Measurement

Why this matters for AI-search measurement

A single rank compresses multiple stages into one number. A more useful measurement stack keeps them separate:

  1. Source inclusion

    Was the source cited at all?

  2. Citation position

    Where did it appear in the observed citation order?

  3. Evidence reflection

    How strongly did its evidence align with the answer?

  4. Downstream outcomes

    Did associated entities or claims appear, and did the answer move into comparison or recommendation?

Rank can remain one useful feature. It should not be treated as a complete measure of narrative impact or answer alignment.

Inference boundary

What this analysis does not prove

“Rank 1” means the first cited-source position in the observed answer. It is not a search-engine ranking and it does not directly reveal the model’s internal retrieval order.

The observed premium could reflect underlying source relevance, the match between evidence and the emerging answer, retrieval or citation-selection characteristics, citation position itself, or a combination of those factors.

  • No position causality

    The study does not show that moving the same source to Rank 1 would cause its alignment share to increase.

  • No internal-order claim

    Observed citation position should not be relabelled as internal retrieval rank.

  • No narrative-control claim

    The small local association does not establish control over whole-answer wording, structure or interpretation.

Causal test

What would establish a citation-order effect?

A stronger experiment would hold substantive information constant and manipulate source order directly:

  1. Fix prompt and source set

    Use the same question and the same evidence in every condition.

  2. Randomize source order

    Move the same source between cited positions rather than comparing naturally different sources.

  3. Repeat generations

    Separate the order effect from ordinary generative variation.

  4. Measure multiple outcomes

    Evaluate source-to-answer alignment, entity and first-mention position, recommendation, factual accuracy and blinded human review where useful.

Conclusion

The narrower conclusion is the stronger one

The whole-answer similarity analysis provides little evidence that a shared Rank-1 source makes separate responses converge. The within-response analysis shows something subtler: first cited evidence received a small alignment premium relative to the opportunities available inside the same answer.

Those findings coexist because they operate at different levels. Rank 1 is associated with a modest local signal—not narrative control.

Research details

How the two analyses were constructed

For the full study design, dataset and four-outcome framework, see Article 1’s research details. Article 2 focuses on the methods needed to interpret similarity and alignment independently.

Pair construction
Whole responses were paired when they shared a source, with a Rank-1 group and a disjoint Rank-3-or-lower comparison group.
Cell weighting
Source/query cells received equal weight to prevent common sources from dominating through overlapping pairs.
Whole-answer measure
Cosine similarity summarized semantic resemblance between each complete response pair.
Evidence processing
Source evidence was chunked, responses were sectioned, and strong chunk-to-section similarities were summarized by source.
Alignment null
Exact source ranks were shuffled within each response, preserving its evidence opportunities and semantic scores.
Eligible population
Responses required usable Rank-1 and lower-rank evidence plus sufficient chunks and response sections for comparison.

Boundaries

Limitations

  • Observational position

    Citation order was observed rather than randomized, so the alignment mechanism is not identified.

  • Selected opportunity set

    The alignment estimate applies to 660 eligible responses, not to answers without usable competing evidence.

  • Semantic proxies

    Cosine similarity and chunk-to-section alignment approximate semantic relationships; they are not direct measures of model reasoning or human-perceived influence.

  • Finance scope

    Results come from finance responses and may differ across domains, platforms or collection conditions.

Narrative Fidelity series

Article 2 establishes the semantic-alignment layer

Continue through the Kojable Research library. All five Narrative Fidelity articles are now published.

FAQ

Frequently asked questions

Do shared Rank-1 sources make AI answers converge?

The study did not find meaningful convergence. Mean whole-answer similarity was 0.7247 for pairs sharing a Rank-1 source and 0.7213 for the lower-rank comparison, a 0.0035 difference whose interval crossed zero.

How large was the Rank-1 semantic-alignment premium?

Among 660 eligible responses, Rank-1 alignment share was 0.4043 versus a 0.3896 shuffled expectation: a modest 0.0147 difference, or 1.47 percentage points.

What does Rank 1 mean in this study?

Rank 1 means the first cited-source position in the observed answer. It is not a search-engine ranking and does not directly reveal internal retrieval order.

Does a 40.4% share mean Rank 1 gets 40% of every answer?

No. It is the average alignment share among eligible responses with usable competing evidence. The valid comparison is the response-specific shuffled expectation of about 39.0%, not a 25% four-bin baseline.

Would changing citation order cause alignment to change?

This observational study cannot establish that. A causal test would hold the prompt and source set fixed, randomize source order and compare repeated generations.

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish