Kojable research · Narrative fidelity · Citation rank

Do Top-Ranked Citations Shape AI-Generated Finance Answers?

Published Updated By Piush Vaish

A study of 1,500 generated finance responses separates four possible rank effects: shared-source similarity, within-response semantic alignment, entity visibility and explicit recommendation.

Key finding

Rank 1 showed a modest semantic-alignment advantage and a clearer entity-visibility association, while shared-source similarity and recommendation effects were not credibly different from zero.

Qualification: This observational study does not establish that changing citation order would cause any of these outcomes.

  • Narrative fidelity
  • Citation rank
  • Finance AI answers
  • 15 min read
Four Rank-1 effect estimates with 95% intervals: semantic alignment and entity mention are positive, while shared-source similarity and recommendation intervals cross zero.
Figure 1. Rank 1 has positive estimated effects for opportunity-aware alignment and adjusted entity mentions. The intervals for shared-source response similarity and adjusted recommendation cross zero, meaning the study remains compatible with no effect in either direction for those outcomes.
Open full-resolution figure
1,500generated finance responses
+1.47ppRank-1 alignment versus shuffled expectation
11.7%adjusted Rank-1 entity mention probability
25 eventspositive recommendations in the eligible analysisLow precisionToo few to support a strong recommendation conclusion.

Study overview

Executive summary

The order of citations in an AI-generated answer feels consequential. A source listed first appears to have won the retrieval contest, making it tempting to assume that its ideas, entity names and recommendations will dominate the response.

The evidence is more restrained. Across 1,500 generated finance responses, Rank-1 evidence had a small but statistically discernible advantage in within-response semantic alignment. Entities associated with Rank-1 sources were also more likely to appear and tended to be introduced earlier.

Two stronger interpretations were not supported. Responses sharing a Rank-1 source were not meaningfully more similar to one another, and higher citation rank did not show a credible increase in explicit recommendation.

Rank 1 carries a modest narrative and visibility advantage, but it is neither narrative control nor endorsement.

Answer first

Do top-ranked citations shape AI-generated finance answers?

  • Modestly in alignment and visibility; not as control or endorsement

    Rank 1 was associated with a 1.47-percentage-point alignment advantage over shuffled ranks and greater entity visibility. The study did not detect meaningful shared-source answer similarity or a reliable recommendation advantage, and its observational design cannot isolate citation position as the cause.

The four outcomes answer different questions. In this study, similarity ≠ alignment ≠ visibility ≠ recommendation.

Study design

What we studied

The dataset contains 1,500 generated responses split evenly across cash flow, payment processing and fraud detection. Prompts covered informational, educational, commercial and transactional intent.

Response sample
500 cash-flow, 500 payment-processing and 500 fraud-detection responses.
Citation coverage
1,171 responses contained at least one cited source; 329 did not.
Citation records
3,797 citation rows across 3,039 distinct canonical sources.
Four outcomes
Shared-source response similarity, opportunity-aware semantic alignment, entity visibility and explicit recommendation.

Separating the outcomes prevents a citation from being treated as a generic unit of “influence.” Each measure represents a different stage between source inclusion and the final wording of an answer.

Finding 1 · Unsupported effect

Shared Rank-1 evidence did not make answers meaningfully more alike

If Rank-1 evidence strongly determined answer construction, responses sharing the same Rank-1 source should be noticeably more similar than responses sharing a lower-ranked source. That pattern did not appear.

Mean cosine similarity was approximately 0.725 for response pairs sharing a Rank-1 source and 0.721 for pairs sharing a disjoint Rank-3-or-lower source. The estimated difference was 0.0035, with a 95% interval from −0.0199 to 0.0253. The directional permutation p-value was 0.484, and the standardised effect was approximately 0.045.

Interpretation: The interval is compatible with a small negative effect, no effect or a small positive effect. The study does not support the claim that sharing a Rank-1 source makes AI answers substantially more similar.

Finding 2 · Supported, but modest

Rank 1 had a modest within-response alignment advantage

A different question produced a different result. Rather than comparing separate answers, this analysis measured which available citation-rank bins aligned most closely with sections of each response. It then shuffled exact source ranks within each response to create the expected alignment share under the null.

Among 660 responses with usable Rank-1 evidence and at least one lower-rank opportunity, the observed Rank-1 alignment share was 0.404. The shuffled expectation was approximately 0.390. The difference was 0.0147, or 1.47 percentage points, with a 95% response-level interval from 0.0052 to 0.0241. The directional permutation p-value was 0.001; the two-sided value was 0.004.

Two-panel comparison showing nearly equal response similarity for shared Rank-1 and lower-rank sources, alongside Rank-1 alignment share modestly above its shuffled expectation.
Figure 2. Sharing a Rank-1 source produced little evidence of greater answer-to-answer similarity, while Rank 1 received a modest within-response alignment premium among eligible answers. These are different estimands and should not be collapsed into one claim about citation influence.
Open full-resolution figure

The size matters. A 1.47-percentage-point advantage is statistically discernible, but it is not dominance. The observational design also does not establish that experimentally moving a source into Rank 1 would make the answer align more closely with it.

Finding 3 · Observational association

Higher-ranked source entities were more visible

Raw entity mention rates declined with citation rank: 12.0% at Rank 1, 10.0% at Rank 2, 8.9% at Ranks 3–5 and 5.9% at Rank 6 or below.

After standardising for evidence brand density, finance topic, query intent and base query, the probabilities remained ordered: 11.7%, 10.2%, 9.0% and 6.5%, respectively. The adjusted Rank-1 minus Ranks-3–5 difference was +2.7 percentage points, with a 95% interval from 0.7 to 5.0. Rank 1 versus Rank 6+ was +5.2 points, with a 95% interval from 2.7 to 8.1. The Rank-1-versus-Rank-2 difference was less certain and should not be treated as established.

Mentioned Rank-1 entities first appeared about 33% of the way through a response, compared with 41% for Rank 2, 44% for Ranks 3–5 and 51% for Rank 6+.

Entity mention probabilities fall across citation-rank groups, while average first-mention position moves later from Rank 1 through Rank 6 and below.
Figure 3. Entity mention rates decline as citation rank falls, and mentioned lower-ranked entities tend to appear later. This is an observational visibility association, not proof that ranking first causes an entity or brand to be mentioned.
Open full-resolution figure

Higher rank may reflect better query fit, more salient entity cues, retrieval quality or other characteristics not fully observed here. Adjustment reduces several obvious alternatives but cannot isolate citation position as the mechanism.

Finding 4 · Unsupported effect

Citation rank did not translate into a credible recommendation advantage

Visibility is not endorsement.

The recommendation analysis covered 1,576 cited entity candidates in commercial and transactional responses, but only 25 positive recommendation events occurred. Raw recommendation rates were 2.03% at Rank 1, 1.63% at Rank 2, 1.62% at Ranks 3–5 and 0% at Rank 6+.

Standardised probabilities were 2.13%, 1.80%, 1.64% and 0.69%, respectively. The adjusted Rank-1 minus Ranks-3–5 difference was +0.48 percentage points, with a 95% interval from −0.90 to 1.78. Fisher’s exact p-value was 0.651.

Low raw and standardised recommendation rates across citation-rank groups with overlapping uncertainty intervals and only 25 positive events overall.
Figure 4. Recommendation events were rare across all rank groups. Overlapping uncertainty and only 25 positive events mean the data do not establish a reliable Rank-1 recommendation advantage.
Open full-resolution figure

Interpretation: The available evidence is unsupported and imprecise, not proof that every small recommendation effect is impossible. It does rule out confidence in citation rank alone as a strong or reliable recommendation mechanism.

Coverage and opportunity

Each outcome has a different eligible denominator

Of 1,500 responses, 1,171 contained at least one cited source and 329 did not. Later analyses required additional opportunities: the alignment analysis needed usable Rank-1 evidence and at least one lower-rank opportunity, while recommendation was limited to cited entity candidates in commercial and transactional answers.

Coverage diagram showing 1,500 total responses, 1,171 with cited sources, 329 without citations, and narrower eligibility sets for alignment, entity and recommendation analyses.
Figure 5. The hypotheses target different populations and opportunity sets. Coverage narrows from all generated responses to cited responses and then to outcome-specific eligible records, so one result should not be generalised to another denominator.
Open full-resolution figure

How far these findings travel

Rank is observed alongside relevance, retrieval quality and source characteristics

Citation position may be associated with source relevance, retrieval quality, evidence salience or other unobserved factors. The design therefore describes associations within this finance response dataset; it cannot isolate citation position as the causal mechanism.

The entity analysis uses domain-derived entity names rather than a fully human-curated brand taxonomy. This provides a consistent, conservative rule, but it can miss ambiguous or non-domain brand forms. Recommendation findings apply only to commercial and transactional responses with cited entity candidates.

Measurement system

Measure the path from inclusion to recommendation in stages

  1. 1. Source inclusion

    Was the source present in the answer’s citation set?

  2. 2. Citation position

    Where did the source appear among the citations?

  3. 3. Semantic and entity reflection

    Did the answer reflect the evidence or mention its associated entity, and how early?

  4. 4. Explicit recommendation

    Did the answer endorse the entity as suitable for the user’s goal?

The study finds evidence of movement across some of the first three stages. It does not establish a credible effect at the fourth.

Bounded interpretation

What the evidence supports

  • A small semantic-alignment advantage

    Top citation position is associated with a modest increase in evidence-to-answer alignment among eligible responses.

  • A clearer entity-visibility association

    Higher-ranked source entities appear more often and earlier, particularly when Rank 1 is compared with sources below Rank 2.

  • Recommendation remains a separate outcome

    Citation rank alone is not reliable evidence that an entity will be recommended.

Claim boundary

What the evidence does not support

  • No narrative control

    Rank 1 did not make answers sharing a source meaningfully more alike and did not absorb most semantic alignment.

  • No endorsement inference

    A first-position citation is not evidence that the AI system endorses or will recommend its associated entity.

  • No position causality

    The study does not show that moving a citation to Rank 1 would cause greater alignment, visibility or recommendation.

Next causal research step

Randomise source order while holding the prompt and evidence fixed

A stronger future experiment should distinguish citation order from source relevance and retrieval quality. It should:

  1. Hold the prompt and source set fixed

    Randomly vary only source ordering across otherwise comparable conditions.

  2. Use repeated generations

    Separate order effects from ordinary model variability.

  3. Add blinded human review where appropriate

    Keep reviewers unaware of source-order assignment when evaluating qualitative outcomes.

  4. Measure the complete outcome sequence

    Record semantic alignment, entity mentions, first-mention position, recommendation and factual accuracy.

This is the proposed causal follow-up, not the design used for the observational study reported on this page.

Study summary

Four outcomes, four separate conclusions

Rank-1 comparisons, inferential conclusions and key reasons.
Outcome Effect Conclusion Reason
Shared-source response similarity +0.0035 versus Rank 3+ Not supported 95% interval crosses zero.
Opportunity-aware alignment share +0.0147 versus shuffled ranks Supported, but modest Positive response-level interval; small magnitude.
Adjusted entity mention probability +0.0271 versus Ranks 3–5 Supported Adjusted interval remains above zero.
Adjusted recommendation probability +0.0048 versus Ranks 3–5 Not supported Interval crosses zero; only 25 positive events.

Methodology and research details

Opportunity-aware tests preserve the question each outcome asks

Shared-source similarity compared response pairs that shared a Rank-1 source with pairs sharing a disjoint Rank-3-or-lower source. Cosine similarity summarised answer-to-answer semantic proximity, with permutation inference and a standardised effect used to assess the observed difference.

Within-response alignment summarised usable evidence at the source level, compared only rank bins present in each eligible answer and shuffled exact source ranks within the response. This preserves the answer’s available citation opportunities while estimating the expected Rank-1 share under exchangeable ranks.

Entity visibility recorded domain-derived entity mentions and first-mention position. Adjusted probabilities standardised over evidence brand density, finance topic, query intent and base query. Recommendation analysis then narrowed the population to cited entity candidates in commercial and transactional responses.

Figures were generated from the study’s retained figure-source data and reproducible build script. The production page serves copied PNG assets so editorial source materials remain separate from public web paths.

Limitations

Interpret position as an association, not an intervention

  • Observational design

    Citation position may travel with relevance, retrieval quality or other unobserved factors. Position was not randomised.

  • Outcome-specific coverage

    329 responses had no cited source, alignment required usable Rank-1 and lower-rank opportunities, and recommendation covered only eligible commercial and transactional responses.

  • Entity taxonomy

    Domain-derived entity names are not equivalent to a fully human-curated brand taxonomy.

  • Rare recommendation events

    Only 25 positive events materially limit precision and the ability to distinguish small differences between rank groups.

  • Finance scope

    Results come from three finance topics and four prompt-intent classes; other domains, models or collection conditions may differ.

Narrative Fidelity series

This flagship synthesis begins the completed five-part sequence

This page establishes the full study and its overall result. All five Narrative Fidelity articles are now published:

Continue through the Kojable Research library or read the related study of query-to-passage semantic relevance. The five-part Narrative Fidelity series is complete.

FAQ

Frequently asked questions

Do top-ranked citations control AI-generated finance answers?

No. Rank 1 had a modest within-response alignment advantage and a clearer entity-visibility association, but shared Rank-1 evidence did not make separate answers meaningfully more alike.

How large was the Rank-1 semantic-alignment advantage?

Among 660 eligible responses, Rank 1 received an alignment share of 0.404 versus an approximately 0.390 shuffled expectation: a 0.0147 difference, or 1.47 percentage points.

Were Rank-1 source entities more visible?

Yes, observationally. Adjusted entity mention probability was 11.7% at Rank 1, compared with 9.0% at Ranks 3–5 and 6.5% at Rank 6+. Mentioned Rank-1 entities also appeared earlier on average.

Did Rank 1 make an entity more likely to be recommended?

The study did not establish a credible recommendation advantage. Only 25 positive events occurred, and the adjusted Rank-1 versus Ranks-3–5 interval crossed zero.

Would moving a source to Rank 1 cause these outcomes?

This observational study cannot answer that causal question. A stronger experiment would hold prompts and source sets fixed, randomise source order and use repeated generations with blinded review where appropriate.

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish