Kojable research · Narrative fidelity · Article 5
What This Study Can—and Cannot—Prove About Citation Rank in AI Answers
Articles 1–4 found different relationships between citation position, semantic reflection, entity visibility and recommendation. Article 5 defines where those findings apply, where they stop, and what experiment is required before claiming citation-position causality.
Key finding
The observational study supports bounded associations with semantic reflection and source-entity visibility, but it does not establish narrative control, endorsement or a causal effect of citation position.
Qualification: Different results apply to different eligible populations, and source position was observed rather than randomized.
- Narrative fidelity
- Research governance
- Causal inference
- Citation rank
- Article 5 of 5
- 18 min read
Series conclusion
Articles 1–4 asked progressively stronger questions
Article 1 established the complete study pattern. Article 2 separated whole-answer similarity from local semantic alignment. Article 3 measured source-entity visibility and first-mention position. Finally, Article 4 tested explicit recommendation.
This capstone asks how far those results can responsibly travel. The answer depends on population, opportunity, outcome definition, uncertainty and the distinction between association and causation.
Direct answer
What does the citation-rank study actually prove?
It supports a modest Rank-1 semantic-alignment association among eligible responses and a clearer association between higher citation position and measured source-entity visibility.
It does not establish whole-answer narrative control, reliable recommendation advantage, internal retrieval priority or a causal effect of moving a source to Rank 1.
-
The strongest next step is a randomized source-order experiment
The same prompt and source set must be held fixed while source position is randomly assigned across repeated generations.
Population governance
One study does not mean one denominator
The full sample contains 1,500 responses: 500 about cash flow, 500 about payment processing and 500 about fraud detection. Prompt intent spans informational, educational, commercial and transactional questions. Not every response can contribute to every hypothesis.
| Population | Count | What it supports |
|---|---|---|
| All generated finance responses | 1,500 | Study-wide coverage |
| Responses with at least one cited source | 1,171 | Observed citation-position analysis |
| Responses with no cited source | 329 | No observed citation position |
| Commercial or transactional responses | 717 | Recommendation context |
| Alignment-opportunity responses | 660 | Primary within-response semantic comparison |
| Recommendation-eligible responses | 542 | Responses with a cited entity candidate |
-
Do not collapse populations
The 660-response alignment conclusion does not describe all 1,500 responses. Recommendation does not describe informational or educational answers, and rank cannot describe the 329 unsourced responses.
Estimand
Opportunity sets are part of the research question
A response with one usable source cannot reveal whether Rank 1 received preferential semantic alignment: there is no lower-ranked competitor. The primary alignment analysis therefore required usable evidence at Rank 1, at least two usable cited sources, at least two usable rank bins, and sufficient evidence and response text for similarity calculations.
-
Hold the response fixed
Compare Rank 1 with lower-ranked evidence inside the same generated answer.
-
Require genuine alternatives
Include only responses where at least two usable ranks can be compared.
-
Shuffle the same ranks
Ask whether observed Rank-1 share exceeds reassignment of the same source ranks within eligible responses.
The resulting 660-response selection is not a defect. It defines the estimand: among responses where Rank 1 and lower-ranked evidence can genuinely be compared, does Rank 1 receive more alignment share than shuffled assignment predicts? It is not evidence that Rank 1 receives more alignment in all AI answers.
Four outcomes
The series measures several outcomes—not one generic “rank effect”
Shared-source similarity, within-response semantic alignment, source-entity visibility and explicit recommendation are different outcomes with different populations and effect scales. Combining them into one “citation influence” metric would erase the study's most important distinctions.
Estimand 2 · Supported and modest
Rank 1 received a small opportunity-aware alignment premium
Across 660 eligible responses, observed Rank-1 alignment share was 0.4043 versus a shuffled expectation of 0.3896. The difference was +0.0147, or +1.47 percentage points, with a 95% interval from +0.0052 to +0.0241. The directional p-value was 0.001 and the two-sided value was 0.004.
-
Semantic association only
Alignment is not copying, factual dependence, narrative dominance, internal retrieval priority or a causal effect of changing source position.
Estimand 3 · Supported association
Higher citation position tracked greater measured entity visibility
The entity analysis used 3,668 unique response–entity observations. Adjusted mention probabilities were 11.7% at Rank 1, 10.2% at Rank 2, 9.0% at Ranks 3–5 and 6.5% at Rank 6+.
Rank 1 vs Ranks 3–5
+2.7 percentage points
95% interval: +0.7 to +5.0 points.
Rank 1 vs Rank 6+
+5.2 percentage points
95% interval: +2.7 to +8.1 points.
For mentioned entities, average first-mention positions were 33.4% at Rank 1, 41.2% at Rank 2, 43.8% at Ranks 3–5 and 50.8% at Rank 6+. Rank 1 versus Rank 2 was less decisive, and most Rank-1 source entities still were not named. The outcome is measured source-entity visibility, not a complete human-curated brand taxonomy.
Estimand 4 · Unsupported and imprecise
Citation position did not yield a credible recommendation advantage
The recommendation analysis contained 1,576 eligible response–entity candidates across 542 eligible responses, with only 25 positive recommendation events. Adjusted probabilities were 2.13% at Rank 1, 1.80% at Rank 2, 1.64% at Ranks 3–5 and 0.69% at Rank 6+.
-
Central contrast: +0.48 percentage points
The 95% interval was −0.90 to +1.78 points and the raw Fisher exact p-value was 0.651. The interval permits small harm or small benefit; it does not prove ranks are equal.
Uncertainty
Confidence intervals carry more information than a binary label
Interpretation should not stop at “significant” or “not significant.” Shared-source similarity permits small negative, zero or small positive effects. Semantic alignment is consistently positive but modest. Entity visibility is positive against materially lower ranks. Recommendation permits small harm or small benefit.
-
Outcome
Name exactly what is measured.
-
Population
Keep the eligible denominator attached.
-
Effect size
Interpret the estimate on its original scale.
-
Uncertainty interval
Ask which substantively different conclusions remain compatible.
Comparability
Adjustment does not create causality
Entity and recommendation probabilities were standardized for evidence entity or brand density, finance topic, query intent and base query. This improves comparability across observed rank groups. It does not randomize citation position.
-
Unmeasured factors remain
Exact query-source relevance, source authority, factual quality, formatting, redundancy, entity-to-claim clarity, citation-selection behavior and latent answer planning may affect both position and outcome.
The standardized results remain observational associations.
Terminology
“Rank 1” is the first observed cited-source position
Rank 1 means the first cited-source position in the observed generated answer.
It is not the first search result, first internally retrieved source, highest internal source score, first document processed by the model or an experimentally assigned first position. The study observes final citation position, not the full hidden retrieval pipeline.
Measurement uncertainty
Domain-derived entity labels require cautious interpretation
Source entities were derived from domains and matched with conservative token-bounded aliases. Publications may differ from companies; companies may operate several domains; product and parent-company names may differ; and abbreviations, descriptive references or ambiguous short names may evade or complicate matching.
Measured source-entity visibility is not equivalent to a complete human-curated brand taxonomy. Future research should validate automated matches against human annotation.
Least mature outcome
Recommendation is rare and linguistically ambiguous
“Consider X,” “X is one option,” “X performs well,” “X may suit some teams” and “choose X” express different strengths of preference. They should not be treated as equivalent endorsements.
A reproducible rule can screen at scale, but stronger future claims require a human-reviewed recommendation taxonomy. With 25 events, this study can reject an overconfident endorsement inference; it cannot precisely estimate subtle differences by rank or entity.
Causal boundary
The study cannot identify what moving a source would cause
The observational findings do not show that moving a source from Rank 3 to Rank 1 would increase semantic alignment, entity mention probability, earlier placement or recommendation probability. Citation position may be jointly determined with relevance, quality, salience and source-set composition.
Counterfactual question
For the same prompt and the same source set, what changes when source position changes and everything else is held constant?
The current study did not randomize source order and cannot answer that question directly.
Next research design
A randomized source-order experiment is needed next
-
Experimental unit
Use prompt × fixed source set across multiple finance topics, query intents, source types and models.
-
Randomization
Hold source content fixed, randomly assign order and balance the same source across several positions within the same prompt and source set.
-
Repeated generation
Generate multiple answers for each assignment so one stochastic response cannot determine the result.
-
Blinded review
Reviewers who do not know assigned order should label source support, semantic reflection, entity mention, first position, framing, shortlist inclusion, recommendation, accuracy and completeness.
-
Within-unit analysis
Estimate within-prompt and within-source effects, separating identity, relevance, position and model variability. Report intervals and heterogeneity by topic, intent, source type and model.
This is the design that could support a causal claim about citation position.
Practical governance
Six rules for using the current evidence
-
1. Treat citation rank as a signal, not a causal lever
Observed associations do not reveal what an intervention would change.
-
2. Keep each outcome separate
Citation, semantic reflection, entity presence, narrative position, positive framing and recommendation are different.
-
3. Attach every metric to its denominator
Keep 1,500 total responses, 1,171 sourced responses, 660 alignment-eligible responses, 3,668 entity observations and 1,576 recommendation candidates distinct.
-
4. Report effect sizes and intervals
P-values do not replace the original scale or uncertainty range.
-
5. Preserve null findings
Similarity and recommendation prevent a cleaner but unsupported narrative-control story.
-
6. Separate diagnosis from intervention
The study diagnoses associations; it does not show that deliberately changing position will change an outcome.
Five-article synthesis
Supported, not supported and not tested causally
Supported
Bounded observational associations
-
Modest Rank-1 within-response semantic-alignment association among eligible responses.
-
Higher citation position associated with greater measured source-entity visibility.
-
Mentioned higher-ranked entities tended to appear earlier.
-
Strongest visibility contrasts occurred between Rank 1 and materially lower positions.
Not supported
Broader rank narratives
-
Shared Rank-1 sources make complete answers meaningfully more similar.
-
Rank 1 controls the narrative or differs clearly from Rank 2 on every visibility outcome.
-
Higher citation rank reliably produces recommendation or equals endorsement.
Not tested causally
Moving the same source
-
Causes more semantic reflection.
-
Causes entity mention or earlier placement.
-
Causes recommendation.
These distinctions are not footnotes. They are the proper interpretation of the research.
Research boundary in one sentence
Association with reflection and visibility—not control or causation
The study supports a modest association between first cited-source position and what an AI-generated finance response reflects or visibly mentions; it does not establish narrative control, endorsement, or a causal effect of citation position.
The next step is not to make the observational claim louder. It is to run the randomized experiment that can test the causal question.
Research details
A governance reference for the complete study
- Study sample
- 1,500 responses across cash flow, payment processing and fraud detection; 500 per topic.
- Prompt intents
- Informational, educational, commercial and transactional.
- Citation coverage
- 1,171 sourced responses and 329 unsourced responses.
- Alignment opportunity
- 660 responses with usable Rank-1 and lower-rank evidence.
- Entity population
- 3,668 unique response–entity observations.
- Recommendation population
- 1,576 candidates across 542 eligible responses; 25 positive events.
- Four estimands
- Whole-answer similarity, within-response alignment, entity visibility and recommendation.
- Design
- Observational analyses with permutation, bootstrap uncertainty and covariate standardization where appropriate.
Detailed methods and outcome-specific diagnostics remain in Articles 1–4. This article governs how their conclusions should be combined and communicated.
Limitations
Each limitation constrains a specific claim
-
Observed, not randomized
Citation position cannot identify an intervention effect because relevance and source composition may be jointly determined with rank.
-
Different eligible populations
Alignment applies to a selected opportunity population; recommendation applies only to eligible commercial and transactional contexts.
-
329 unsourced responses
Citation-position conclusions cannot directly describe answers with no observed citation.
-
Entity identity uncertainty
Domain-derived labels and automated aliases can miss or imperfectly classify brand references.
-
Sparse and ambiguous recommendation
Twenty-five positive events and varied recommendation language limit precision and endorsement claims.
-
Unmeasured confounding and model variability
Adjustment cannot capture every relevance, quality, planning or generation factor.
-
Finance-domain scope
Results may not generalize to other domains, models, platforms or collection conditions.
-
Hidden retrieval pipeline
Final citation position does not directly expose internal retrieval, scoring or processing order.
Complete Narrative Fidelity series
All five research articles are now published
-
Article 1 — Flagship synthesis
Do Top-Ranked Citations Shape AI-Generated Finance Answers?
-
Article 2 — Semantic alignment
Citation Rank and Semantic Alignment: Why Shared Sources Don’t Make AI Answers Converge
-
Article 3 — Brand visibility
Citation Rank and Brand Visibility: Why Higher-Ranked Source Entities Appear More Often and Earlier
-
Article 4 — Recommendation
A Citation Is Not an Endorsement: Why Higher Citation Rank Did Not Translate Into Recommendation
-
Article 5 — Research governance
What This Study Can—and Cannot—Prove About Citation Rank in AI Answers
This article completes the five-part Narrative Fidelity series. Continue through the Kojable Research library.
FAQ
Frequently asked questions
Does this study prove that moving a source to Rank 1 changes an AI answer?
No. Citation position was observed rather than randomized, so the study cannot identify the causal effect of changing source position.
Why do the analyses use different numbers of responses?
Each outcome requires a different eligible population. For example, semantic alignment requires usable Rank-1 and lower-rank evidence, while recommendation is evaluated only in eligible commercial and transactional contexts.
Does statistical adjustment make the entity-visibility result causal?
No. Adjustment improves comparability for measured factors but cannot eliminate all unmeasured confounding.
What does Rank 1 mean in this research?
Rank 1 means the first cited-source position in the observed generated answer. It does not mean first search result or directly observed internal retrieval order.
What experiment is needed next?
A randomized source-order experiment holding the prompt and source set fixed, with repeated generations and blinded outcome review.