Kojable research · Narrative fidelity · Article 5

What This Study Can—and Cannot—Prove About Citation Rank in AI Answers

Published By Piush Vaish

Articles 1–4 found different relationships between citation position, semantic reflection, entity visibility and recommendation. Article 5 defines where those findings apply, where they stop, and what experiment is required before claiming citation-position causality.

Key finding

The observational study supports bounded associations with semantic reflection and source-entity visibility, but it does not establish narrative control, endorsement or a causal effect of citation position.

Qualification: Different results apply to different eligible populations, and source position was observed rather than randomized.

  • Narrative fidelity
  • Research governance
  • Causal inference
  • Citation rank
  • Article 5 of 5
  • 18 min read
Coverage and eligibility counts showing that the 1,500-response study produces different eligible populations for citation, semantic-alignment and recommendation analyses.
Figure 1. The study begins with 1,500 generated finance responses, but each question has its own eligible population. A result remains interpretable only while its denominator and opportunity set stay attached to it.
Open full-resolution figure
1,500total generated finance responses
1,171responses with a cited source
660semantic-alignment-eligible responses
25positive recommendation eventsLimited precision

Series conclusion

Articles 1–4 asked progressively stronger questions

Article 1 established the complete study pattern. Article 2 separated whole-answer similarity from local semantic alignment. Article 3 measured source-entity visibility and first-mention position. Finally, Article 4 tested explicit recommendation.

This capstone asks how far those results can responsibly travel. The answer depends on population, opportunity, outcome definition, uncertainty and the distinction between association and causation.

Direct answer

What does the citation-rank study actually prove?

It supports a modest Rank-1 semantic-alignment association among eligible responses and a clearer association between higher citation position and measured source-entity visibility.

It does not establish whole-answer narrative control, reliable recommendation advantage, internal retrieval priority or a causal effect of moving a source to Rank 1.

  • The strongest next step is a randomized source-order experiment

    The same prompt and source set must be held fixed while source position is randomly assigned across repeated generations.

Population governance

One study does not mean one denominator

The full sample contains 1,500 responses: 500 about cash flow, 500 about payment processing and 500 about fraud detection. Prompt intent spans informational, educational, commercial and transactional questions. Not every response can contribute to every hypothesis.

Each research question has a distinct eligible population
Population Count What it supports
All generated finance responses 1,500 Study-wide coverage
Responses with at least one cited source 1,171 Observed citation-position analysis
Responses with no cited source 329 No observed citation position
Commercial or transactional responses 717 Recommendation context
Alignment-opportunity responses 660 Primary within-response semantic comparison
Recommendation-eligible responses 542 Responses with a cited entity candidate
  • Do not collapse populations

    The 660-response alignment conclusion does not describe all 1,500 responses. Recommendation does not describe informational or educational answers, and rank cannot describe the 329 unsourced responses.

Estimand

Opportunity sets are part of the research question

A response with one usable source cannot reveal whether Rank 1 received preferential semantic alignment: there is no lower-ranked competitor. The primary alignment analysis therefore required usable evidence at Rank 1, at least two usable cited sources, at least two usable rank bins, and sufficient evidence and response text for similarity calculations.

  1. Hold the response fixed

    Compare Rank 1 with lower-ranked evidence inside the same generated answer.

  2. Require genuine alternatives

    Include only responses where at least two usable ranks can be compared.

  3. Shuffle the same ranks

    Ask whether observed Rank-1 share exceeds reassignment of the same source ranks within eligible responses.

The resulting 660-response selection is not a defect. It defines the estimand: among responses where Rank 1 and lower-ranked evidence can genuinely be compared, does Rank 1 receive more alignment share than shuffled assignment predicts? It is not evidence that Rank 1 receives more alignment in all AI answers.

Four outcomes

The series measures several outcomes—not one generic “rank effect”

Shared-source similarity, within-response semantic alignment, source-entity visibility and explicit recommendation are different outcomes with different populations and effect scales. Combining them into one “citation influence” metric would erase the study's most important distinctions.

Four Rank-1 contrasts showing positive semantic-alignment and entity-visibility estimates, while shared-source similarity and recommendation intervals cross zero.
Figure 2. Alignment and entity visibility have positive supported intervals; shared-source similarity and recommendation cross zero. Each estimate describes a different outcome and population, not one overall Rank-1 effect.
Open full-resolution figure

Estimand 1 · Unsupported

Shared Rank-1 sources did not make complete answers converge

Rank-1 mean similarity
0.7247
Lower-rank comparison
0.7213
Difference
+0.0035
95% interval
−0.0199 to +0.0253

The directional permutation p-value was 0.484. The interval is compatible with a small negative, zero or small positive whole-answer effect. This is a convergence test; it does not measure source-to-answer semantic reflection or factual dependence.

Estimand 2 · Supported and modest

Rank 1 received a small opportunity-aware alignment premium

Across 660 eligible responses, observed Rank-1 alignment share was 0.4043 versus a shuffled expectation of 0.3896. The difference was +0.0147, or +1.47 percentage points, with a 95% interval from +0.0052 to +0.0241. The directional p-value was 0.001 and the two-sided value was 0.004.

  • Semantic association only

    Alignment is not copying, factual dependence, narrative dominance, internal retrieval priority or a causal effect of changing source position.

Estimand 3 · Supported association

Higher citation position tracked greater measured entity visibility

The entity analysis used 3,668 unique response–entity observations. Adjusted mention probabilities were 11.7% at Rank 1, 10.2% at Rank 2, 9.0% at Ranks 3–5 and 6.5% at Rank 6+.

Rank 1 vs Ranks 3–5

+2.7 percentage points

95% interval: +0.7 to +5.0 points.

Rank 1 vs Rank 6+

+5.2 percentage points

95% interval: +2.7 to +8.1 points.

For mentioned entities, average first-mention positions were 33.4% at Rank 1, 41.2% at Rank 2, 43.8% at Ranks 3–5 and 50.8% at Rank 6+. Rank 1 versus Rank 2 was less decisive, and most Rank-1 source entities still were not named. The outcome is measured source-entity visibility, not a complete human-curated brand taxonomy.

Estimand 4 · Unsupported and imprecise

Citation position did not yield a credible recommendation advantage

The recommendation analysis contained 1,576 eligible response–entity candidates across 542 eligible responses, with only 25 positive recommendation events. Adjusted probabilities were 2.13% at Rank 1, 1.80% at Rank 2, 1.64% at Ranks 3–5 and 0.69% at Rank 6+.

  • Central contrast: +0.48 percentage points

    The 95% interval was −0.90 to +1.78 points and the raw Fisher exact p-value was 0.651. The interval permits small harm or small benefit; it does not prove ranks are equal.

Uncertainty

Confidence intervals carry more information than a binary label

Interpretation should not stop at “significant” or “not significant.” Shared-source similarity permits small negative, zero or small positive effects. Semantic alignment is consistently positive but modest. Entity visibility is positive against materially lower ranks. Recommendation permits small harm or small benefit.

  1. Outcome

    Name exactly what is measured.

  2. Population

    Keep the eligible denominator attached.

  3. Effect size

    Interpret the estimate on its original scale.

  4. Uncertainty interval

    Ask which substantively different conclusions remain compatible.

Comparability

Adjustment does not create causality

Entity and recommendation probabilities were standardized for evidence entity or brand density, finance topic, query intent and base query. This improves comparability across observed rank groups. It does not randomize citation position.

  • Unmeasured factors remain

    Exact query-source relevance, source authority, factual quality, formatting, redundancy, entity-to-claim clarity, citation-selection behavior and latent answer planning may affect both position and outcome.

The standardized results remain observational associations.

Terminology

“Rank 1” is the first observed cited-source position

Rank 1 means the first cited-source position in the observed generated answer.

It is not the first search result, first internally retrieved source, highest internal source score, first document processed by the model or an experimentally assigned first position. The study observes final citation position, not the full hidden retrieval pipeline.

Measurement uncertainty

Domain-derived entity labels require cautious interpretation

Source entities were derived from domains and matched with conservative token-bounded aliases. Publications may differ from companies; companies may operate several domains; product and parent-company names may differ; and abbreviations, descriptive references or ambiguous short names may evade or complicate matching.

Measured source-entity visibility is not equivalent to a complete human-curated brand taxonomy. Future research should validate automated matches against human annotation.

Least mature outcome

Recommendation is rare and linguistically ambiguous

“Consider X,” “X is one option,” “X performs well,” “X may suit some teams” and “choose X” express different strengths of preference. They should not be treated as equivalent endorsements.

A reproducible rule can screen at scale, but stronger future claims require a human-reviewed recommendation taxonomy. With 25 events, this study can reject an overconfident endorsement inference; it cannot precisely estimate subtle differences by rank or entity.

Causal boundary

The study cannot identify what moving a source would cause

The observational findings do not show that moving a source from Rank 3 to Rank 1 would increase semantic alignment, entity mention probability, earlier placement or recommendation probability. Citation position may be jointly determined with relevance, quality, salience and source-set composition.

Counterfactual question

For the same prompt and the same source set, what changes when source position changes and everything else is held constant?

The current study did not randomize source order and cannot answer that question directly.

Next research design

A randomized source-order experiment is needed next

  1. Experimental unit

    Use prompt × fixed source set across multiple finance topics, query intents, source types and models.

  2. Randomization

    Hold source content fixed, randomly assign order and balance the same source across several positions within the same prompt and source set.

  3. Repeated generation

    Generate multiple answers for each assignment so one stochastic response cannot determine the result.

  4. Blinded review

    Reviewers who do not know assigned order should label source support, semantic reflection, entity mention, first position, framing, shortlist inclusion, recommendation, accuracy and completeness.

  5. Within-unit analysis

    Estimate within-prompt and within-source effects, separating identity, relevance, position and model variability. Report intervals and heterogeneity by topic, intent, source type and model.

This is the design that could support a causal claim about citation position.

Practical governance

Six rules for using the current evidence

  1. 1. Treat citation rank as a signal, not a causal lever

    Observed associations do not reveal what an intervention would change.

  2. 2. Keep each outcome separate

    Citation, semantic reflection, entity presence, narrative position, positive framing and recommendation are different.

  3. 3. Attach every metric to its denominator

    Keep 1,500 total responses, 1,171 sourced responses, 660 alignment-eligible responses, 3,668 entity observations and 1,576 recommendation candidates distinct.

  4. 4. Report effect sizes and intervals

    P-values do not replace the original scale or uncertainty range.

  5. 5. Preserve null findings

    Similarity and recommendation prevent a cleaner but unsupported narrative-control story.

  6. 6. Separate diagnosis from intervention

    The study diagnoses associations; it does not show that deliberately changing position will change an outcome.

Five-article synthesis

Supported, not supported and not tested causally

Supported

Bounded observational associations

  • Modest Rank-1 within-response semantic-alignment association among eligible responses.

  • Higher citation position associated with greater measured source-entity visibility.

  • Mentioned higher-ranked entities tended to appear earlier.

  • Strongest visibility contrasts occurred between Rank 1 and materially lower positions.

Not supported

Broader rank narratives

  • Shared Rank-1 sources make complete answers meaningfully more similar.

  • Rank 1 controls the narrative or differs clearly from Rank 2 on every visibility outcome.

  • Higher citation rank reliably produces recommendation or equals endorsement.

Not tested causally

Moving the same source

  • Causes more semantic reflection.

  • Causes entity mention or earlier placement.

  • Causes recommendation.

These distinctions are not footnotes. They are the proper interpretation of the research.

Research boundary in one sentence

Association with reflection and visibility—not control or causation

The study supports a modest association between first cited-source position and what an AI-generated finance response reflects or visibly mentions; it does not establish narrative control, endorsement, or a causal effect of citation position.

The next step is not to make the observational claim louder. It is to run the randomized experiment that can test the causal question.

Research details

A governance reference for the complete study

Study sample
1,500 responses across cash flow, payment processing and fraud detection; 500 per topic.
Prompt intents
Informational, educational, commercial and transactional.
Citation coverage
1,171 sourced responses and 329 unsourced responses.
Alignment opportunity
660 responses with usable Rank-1 and lower-rank evidence.
Entity population
3,668 unique response–entity observations.
Recommendation population
1,576 candidates across 542 eligible responses; 25 positive events.
Four estimands
Whole-answer similarity, within-response alignment, entity visibility and recommendation.
Design
Observational analyses with permutation, bootstrap uncertainty and covariate standardization where appropriate.

Detailed methods and outcome-specific diagnostics remain in Articles 1–4. This article governs how their conclusions should be combined and communicated.

Limitations

Each limitation constrains a specific claim

  • Observed, not randomized

    Citation position cannot identify an intervention effect because relevance and source composition may be jointly determined with rank.

  • Different eligible populations

    Alignment applies to a selected opportunity population; recommendation applies only to eligible commercial and transactional contexts.

  • 329 unsourced responses

    Citation-position conclusions cannot directly describe answers with no observed citation.

  • Entity identity uncertainty

    Domain-derived labels and automated aliases can miss or imperfectly classify brand references.

  • Sparse and ambiguous recommendation

    Twenty-five positive events and varied recommendation language limit precision and endorsement claims.

  • Unmeasured confounding and model variability

    Adjustment cannot capture every relevance, quality, planning or generation factor.

  • Finance-domain scope

    Results may not generalize to other domains, models, platforms or collection conditions.

  • Hidden retrieval pipeline

    Final citation position does not directly expose internal retrieval, scoring or processing order.

Complete Narrative Fidelity series

All five research articles are now published

This article completes the five-part Narrative Fidelity series. Continue through the Kojable Research library.

FAQ

Frequently asked questions

Does this study prove that moving a source to Rank 1 changes an AI answer?

No. Citation position was observed rather than randomized, so the study cannot identify the causal effect of changing source position.

Why do the analyses use different numbers of responses?

Each outcome requires a different eligible population. For example, semantic alignment requires usable Rank-1 and lower-rank evidence, while recommendation is evaluated only in eligible commercial and transactional contexts.

Does statistical adjustment make the entity-visibility result causal?

No. Adjustment improves comparability for measured factors but cannot eliminate all unmeasured confounding.

What does Rank 1 mean in this research?

Rank 1 means the first cited-source position in the observed generated answer. It does not mean first search result or directly observed internal retrieval order.

What experiment is needed next?

A randomized source-order experiment holding the prompt and source set fixed, with repeated generations and blinded outcome review.

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish