Kojable research · Decision 7 · Finalist source quality

Do Final AI Citation Sources Show Evidence of Selective Quality Filtering?

Published By Piush Vaish

Final citation survivors converge across AI platforms and persist across repeated questions, but those downstream patterns cannot reveal whether authority, passage quality or another hidden criterion caused a candidate to survive.

Key finding

Same-prompt final source sets are about three times as similar as matched controls, while repeated prompts retain substantially more stable source portfolios across adjacent periods.

Qualification: These are downstream reproducibility signals—not direct measurements of source authority, passage quality or a candidate-level quality gate.

  • Source quality
  • Cross-platform convergence
  • Temporal persistence
  • 14 min read
Diagram contrasting observable final citation survivor evidence with missing candidate-level evidence such as rejected passages, authority labels, passage-quality labels and a Decision 7 pass-or-fail outcome.
Figure 1. Decision 7 observes structured final citation survivors but lacks the rejected candidate population, independent quality labels and pass-or-fail outcome required to validate the hidden finalist-quality gate.
Open full-resolution figure
~3×higher same-prompt source similarity than matched controls
~0.46 vs ~0.13same-template vs matched-control temporal similarity
3 dimensionspassage support, source authority and passage quality
Gate unprovenrejected candidates and independent quality labels are missingBoundarySurvivor evidence is not candidate-level gate evidence.

Study overview

Executive summary

A source can be relevant to a query and still be a weak choice for the final answer. Decision 7 asks whether a surviving source and passage are strong enough to become a finalist before an answer is assembled.

The historical export cannot directly observe that decision. It contains final citation survivors rather than the rejected candidates, independent quality labels and pass-or-fail outcomes needed to identify a candidate-level gate.

The survivors nevertheless show strong downstream structure. Across platforms, final domain sets for the same prompt have approximately 0.18 similarity, compared with about 0.06 for matched different-prompt controls. The approximate 0.12 difference means same-prompt sets are about three times as similar.

Repeated prompt templates also retain more similar final source portfolios across adjacent periods: approximately 0.46 for the same template versus approximately 0.13 for matched alternatives, an approximate difference of 0.32.

Final citation survivors are query-specific, cross-platform convergent, temporally persistent and semantically aligned in the observable subset. These are strong downstream signals consistent with selective finalist evaluation, but the candidate-level quality gate remains unproven.

Direct answer

Do final AI citation sources show evidence of selective quality filtering?

  • Structured survivor evidence, not direct gate validation

    Final source choices show strong downstream structure consistent with selective evaluation, but the historical export cannot directly validate the hidden candidate-level quality gate because rejected candidates, independent authority labels, passage-quality labels and the Decision 7 pass-or-fail outcome are absent.

Diagnostic sequence

Decision 7 follows the evidence path to finalist quality

  1. Decision 1

    Retrieval need and citation exposure.

  2. Decision 2

    Buyer-question context and citation exposure.

  3. Decision 3

    Ecosystem source availability when citations are absent.

  4. Decision 4

    Selected source-family composition.

  5. Decision 5

    Content-access and extraction observability.

  6. Decision 6

    Passage-level semantic alignment.

  7. Decision 7

    Finalist quality and downstream survivor structure.

Conceptual boundary

What finalist quality means

Finalist quality is more demanding than relevance alone. A passage may match the query while still being weakly sourced, vague, repetitive, outdated, unsupported or too low quality to anchor a final answer.

Passage support
Does the passage directly answer or provide evidence for the buyer question?
Source authority
Is the publisher credible for this specific topic?
Passage quality
Is the actual text clear, specific, factual, coherent, original and useful?

A candidate can be relevant but weakly sourced, authoritative but irrelevant, relevant and authoritative but vague, or strong across all three dimensions. The dimensions should not be collapsed into a single public quality score.

Identification requirement

Why direct proof requires candidate-level data

A direct test would compare candidates within the same retrieval opportunity and ask whether stronger passage support, higher topic-specific authority and stronger passage quality predict survival at the finalist-quality stage.

  • Candidate denominator missing

    The export does not contain every candidate entering the quality stage or rejected candidate passages.

  • Independent labels missing

    There is no externally validated topic-authority label or passage-quality label.

  • Gate outcome missing

    The Decision 7 pass-or-fail result is unavailable and cannot be separated cleanly from the later answer-readiness stage.

Measurement boundary

Survivor evidence is not gate evidence

Survivor evidence ≠ gate evidence. A property of a surviving output is not automatically evidence of the rule that caused it to survive. Final survivors do not reveal why missing candidates disappeared.

A missing source may never have been retrieved, may have failed access or extraction, may have been semantically weak, may have failed finalist quality or answer readiness, or may have been removed for duplication, policy or presentation. The final output cannot distinguish among those histories.

Downstream test 1

Same-prompt source sets converge across platforms

Final domain sets were compared across platforms answering the same prompt and against matched controls drawn from different prompts in similar query and collection contexts.

Bar chart showing approximate final-domain similarity of 0.18 for same-prompt platform comparisons versus 0.06 for matched different-prompt controls.
Figure 2. AI platforms answering the same prompt produce final source-domain sets that are approximately three times as similar as matched different-prompt controls. The result is a downstream consistency signal, not a calibrated measure of authority or quality.
Open full-resolution figure
Approximate final-domain similarity
Comparison Approx. domain-set similarity
Same prompt~0.18
Matched different prompt~0.06

Approximate difference: ~0.12. Same-prompt similarity is about the matched-control level. The absolute overlap remains moderate; the important result is the relative, query-specific difference.

Interpretation boundary

Cross-platform convergence is not authority

Agreement between platforms does not prove truth, credibility, authority, factual accuracy or high passage quality. Platforms may converge because a source is relevant, visible in search, widely linked, frequently discussed, consistently indexed, prominent in the category or genuinely authoritative.

  • Reproducibility, not rank

    Cross-platform convergence is a downstream reproducibility signal, not a calibrated source-quality measure.

Downstream test 2

Selected source sets persist across adjacent periods

The second downstream test asks whether repeated prompt templates retain similar final source-domain sets in the following period.

Bar chart showing approximate adjacent-period source similarity of 0.46 for the same prompt template versus 0.13 for matched different-template controls.
Figure 3. Repeated prompt templates retain substantially more similar final source-domain sets across adjacent periods than matched alternative prompts. Persistence is operationally important but does not prove that persistent sources are higher quality.
Open full-resolution figure
Approximate adjacent-period domain-set similarity
Comparison Approx. adjacent-period similarity
Same prompt template~0.46
Matched different template~0.13

The approximate difference is ~0.32. Repeated questions retain substantially more similar final source portfolios than matched alternative questions.

Interpretation boundary

Temporal persistence is not proof of quality

Stable source selection can reflect stable search rankings, persistent indexing, long-lived category prominence, repeated popularity, platform defaults, slow-changing content or genuinely durable quality. The historical data cannot distinguish these explanations directly.

Final source selection is reproducible across adjacent periods for repeated questions. Repeated citation is not an authority score or proof that a source is high quality.

Supporting context

What Decision 6 contributes

Decision 6 establishes a downstream semantic property for a selected Gemini subset: observable passage-like final citation text aligns more strongly with its own query than with matched alternatives.

That result does not establish source authority, passage quality, finalist quality or the internal rule used to decide which candidate survived. Semantic relevance and finalist quality remain different concepts.

Measurement framework

Passage support, source authority and passage quality

Three-part diagram separating passage support, source authority and passage quality as independent dimensions contributing to finalist-quality assessment.
Figure 4. Finalist quality should separate passage support, topic-specific source authority and passage quality. A candidate can be strong on one dimension and weak on another, so citation frequency or source-family labels should not be treated as a universal quality score.
Open full-resolution figure
Passage support
Direct evidence, a specific answer, claim support and query-specific usefulness.
Source authority
Topic expertise, primary-source status, transparency, editorial standards and evidence track record.
Passage quality
Clarity, specificity, factual density, originality, coherence and usefulness.

Taxonomy boundary

Source-family labels are not authority labels

Broad classes such as brand-owned, comparable vendors, official documentation, community, code repositories, news/media and other third-party publishers describe the type of source that appeared. They do not establish how credible that source was for a specific claim.

Source-family labels describe source type, not calibrated topic-specific authority. A family can contain both strong and weak evidence, so it should not be interpreted as a quality hierarchy.

Evidence separation

Consensus and semantic relevance should not be collapsed

Cross-platform consensus and semantic relevance can coexist, but the within-template evidence does not establish that consensus reliably identifies more semantically relevant passages. Agreement does not equal truth, and relevance does not prove finalist survival.

The public result is qualitative because the small bridge subset cannot support reconstructive precision. No dedicated consensus chart, source-family ranking or precise bridge table is published.

Supported conclusions

What Decision 7 establishes

  • Direct gate unidentifiable

    Rejected candidates, quality labels and the gate outcome are missing.

  • Query-specific selection

    Same-prompt source sets converge much more than matched different-prompt controls.

  • Temporal reproducibility

    Repeated prompt templates retain substantially more similar source portfolios than matched alternatives.

  • Structured survivors

    Final source selections show strong downstream structure across platforms and time.

  • Plausibility, not proof

    Multiple unobserved mechanisms can produce the observed survivor patterns.

Claim boundary

What Decision 7 does not establish

  • No authority inference

    The study does not show that a source survived because it was authoritative or that rejected sources were less authoritative.

  • No quality hierarchy

    Citation frequency, consensus and source-family classes are not authority or quality scores.

  • No causal gate claim

    The evidence does not show that semantic relevance caused finalist survival, that a universal threshold exists, or that Decision 7 rather than the next stage caused citation emission.

Failure to prove the gate is not evidence that the gate is false. The result is best classified as not proven / supported downstream / gate unvalidated.

Identifiability

No survivor-only model can recover the missing gate

If the final answer contains source A but not sources B, C and D, those absent sources could have failed at retrieval, access, relevance, finalist quality, answer readiness, duplication, policy or presentation. Every history can produce the same final record.

A model trained only on survivors would learn which sources tend to appear in final answers. It would not identify which finalist-quality criteria caused a candidate to survive.

Required telemetry

What evidence would directly validate Decision 7?

  1. Retain the complete candidate set

    Record every passage entering the quality stage, including failures.

  2. Measure dimensions separately

    Add blinded labels for passage support, topic-specific authority and passage quality.

  3. Record the gate outcome

    Preserve pre- and post-quality rank, a pass-or-fail outcome, rejection reason and the later answer-readiness result separately.

Record failures as carefully as survivors.

Future research

A stronger confirmatory study

A confirmatory study should sample complete candidate sets rather than more repetitions of final answers. It should retain all candidates entering Decision 7, obtain blinded ratings for the three quality dimensions, record component scores where available, model survival within each retrieval run, control for platform, query, rank, time and model version, and keep later answer-readiness outcomes separate.

The key question is whether each quality dimension adds independent predictive value within the same candidate opportunity. Controlled interventions would be required for causal validation.

AEO and GEO implications

Why finalist quality matters for AI visibility

Presence in the information ecosystem is not enough. Semantic relevance is not enough either. A source competing for inclusion may also need to be credible for the topic, explicit, specific, well supported, clear, useful and durable enough to survive comparison with alternatives.

  • Make important evidence easy to verify and specific enough to stand alone.

  • Show why the source is qualified to make the claim.

  • Prefer concrete support over generic marketing language.

  • Keep evidence current, clear and usable without reconstructing the argument.

These practical quality questions should be tested directly rather than inferred from citation frequency.

Research relationship

Decision 7 separates relevance from finalist quality

Decision 6 shows that selected passage-like Gemini citation text is strongly query-specific. Decision 7 asks whether semantic fit is sufficient. It is not: a semantically aligned passage can still be weakly sourced, vague, poorly written, outdated or unsupported.

The current evidence supports selective final source behaviour downstream, while authority and passage quality remain unmeasured.

Next stage

Answer readiness remains separate

Decision 8 examines whether a surviving finalist can be used cleanly in the final answer—and why rendered answer quality still cannot reveal the hidden passage-level readiness gate without finalist failures and claim-level citation mappings.

Limitations

Interpret the survivor evidence narrowly

  • Survivor-only and observational

    There is no complete candidate set, rejected-candidate population or Decision 7 outcome.

  • Quality labels unavailable

    Authority and passage quality are not independently measured.

  • Alternative mechanisms remain

    Visibility, search ranking, indexing and platform defaults can produce convergence or persistence.

  • Passage evidence is selected

    Passage-level semantic evidence remains limited to the already-public selected Gemini subset.

The study provides strong evidence about the structure of final source selection. It does not directly identify an internal finalist-quality mechanism.

Research governance

Public-safe research details

This publication preserves the substantive downstream findings while removing customer-identifying and reconstructive details. Reporting uses rounded consistency metrics and generic source-family language.

The public page omits the underlying target company, named competitors, domains, exact source-family consensus values and counts, exact platform and prompt samples, exact comparison counts, confidence intervals, significance values, platform shares, bridge statistics, collection dates, internal paths, script names, provider identity and private candidate-review materials.

Structured data follows the same disclosure boundary as the visible article. The diagrams communicate measurement concepts rather than claiming to reveal proprietary platform architecture.

FAQ

Frequently asked questions

Do final AI citation sources show evidence of quality filtering?

Final source choices show strong downstream structure consistent with selective evaluation, but the historical export cannot directly validate the hidden candidate-level quality gate because rejected candidates, independent authority labels, passage-quality labels and the Decision 7 pass-or-fail outcome are absent.

What does Decision 7 actually observe?

It observes final citation domains, final source combinations, cross-platform source overlap, source persistence across adjacent periods, broad source-family classifications and the already-public semantic-alignment result for a selected Gemini fragment subset. It does not observe rejected candidates or the finalist gate itself.

Why does cross-platform source convergence matter?

AI platforms answering the same prompt produce final domain sets with approximately 0.18 similarity, compared with approximately 0.06 for matched different-prompt controls. The roughly threefold difference shows that final source selection is query-specific rather than an arbitrary draw from a broad source pool.

Does cross-platform agreement prove source authority?

No. Platforms can converge because of relevance, visibility, search rankings, popularity, indexing, shared ecosystem structure or genuine authority. Agreement is a reproducibility signal, not a calibrated measure of truth, authority or passage quality.

Why does temporal source persistence matter?

Repeated prompt templates retain final source portfolios with approximately 0.46 adjacent-period similarity, compared with approximately 0.13 for matched different-template controls. That persistence shows reproducible final source selection for repeated questions.

Does repeated citation prove a source is high quality?

No. Persistence can result from stable search rankings, indexing, visibility, platform defaults, long-lived popularity or genuinely durable quality. The historical data cannot separate those explanations, so repeated citation is not an authority or quality score.

What are the three dimensions of finalist quality?

Passage support asks whether the text directly supports the buyer question, source authority asks whether the publisher is credible for that topic, and passage quality asks whether the text is clear, specific, factual, coherent, original and useful. The dimensions should be measured separately.

Why can't final citations reveal the hidden quality gate?

Final citations contain only survivors. Missing sources may never have been retrieved, may have failed access or relevance, may have failed finalist quality or answer readiness, or may have been removed for duplication, policy or presentation. The same final record is compatible with many different histories.

What evidence would directly validate finalist-quality filtering?

A direct study needs the complete candidate set entering the quality stage, exact candidate passages, separate relevance, authority and passage-quality labels, pre- and post-quality ranks, structured rejection reasons, a Decision 7 pass-or-fail outcome and the later Decision 8 outcome kept separate.

How does Decision 7 connect to Decisions 1–6?

Decisions 1–5 move from retrieval need and query context through source availability, selected-source composition and content-access observability. Decision 6 tests semantic alignment in selected passages, while Decision 7 asks whether final survivor structure is consistent with a later finalist-quality stage without claiming that the hidden gate was observed.

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish