Kojable research · Decision 6 · Passage semantic relevance

Do Final AI Citation Passages Match the Meaning of the Query?

Published By Piush Vaish

Selected Gemini citation fragments align much more strongly with their own queries than with plausible matched alternatives, but final-output alignment does not reveal the hidden relevance gate that produced those survivors.

Key finding

Selected citation fragments ranked first against five matched alternatives about 54% of the time under the primary semantic proxy, versus roughly 17% expected by chance.

Qualification: The analysis measures selected final-output fragments, not Gemini's internal candidate set, relevance score, threshold or rejected passages.

  • Passage relevance
  • Matched-decoy analysis
  • Gemini citation fragments
  • 14 min read
Two-panel chart showing an approximate semantic similarity advantage of 0.18 and lexical advantage of 0.04, alongside top-1 rates of about 54% and 51% compared with a roughly 17% chance benchmark.
Figure 1. Selected Gemini citation fragments align more strongly with their own prompt than with matched alternatives under both latent-semantic and lexical representations. Top-1 identification is roughly three times the chance benchmark.
Open full-resolution figure
~6,000usable Gemini prompt-fragment pairs
~0.27 vs ~0.09own-prompt vs matched-decoy semantic similarity
~54% vs ~17%selected-fragment top-1 rate vs chance
5 of 5well-covered buyer-question types with positive alignmentScopeSelected Gemini fragments, not candidate telemetry.

Study overview

Executive summary

A citation can point to a relevant page while still failing to identify text that meaningfully answers the user's question. Decision 6 moves from page-level evidence to a more demanding test: whether passage-like text selected in a final AI citation matches the meaning and information need of the originating query.

For a selected subset of Gemini citations, emitted URLs contain browser text-fragment pointers that identify text on the cited page. We decoded those final-output anchors and compared each fragment with its own prompt and with five contextually plausible matched alternatives.

Across roughly 6,000 usable Gemini prompt-fragment pairs representing around 300 distinct prompt templates, own-prompt similarity was approximately 0.27 under the primary analyst-created latent-semantic representation, compared with approximately 0.09 for matched decoys. The alignment advantage was therefore about 0.18.

The selected fragment ranked first against five matched alternatives in roughly 54% of template-level comparisons, versus approximately 17% expected by chance. A simpler analyst-created lexical representation produced a smaller advantage of about 0.04 and a top-1 rate of roughly 51%, pointing in the same direction.

Selected Gemini citation fragments show strong query-specific semantic alignment, but the platform's internal candidate-level relevance gate remains unproven.

Direct answer

Do passage-like fragments selected in final AI citations match the query that produced them?

  • Yes, for the observable Gemini subset

    Selected fragments are materially more aligned with their own prompts than with plausible matched alternatives. But this final-output result does not prove Gemini's hidden candidate-level semantic-relevance gate, threshold, reranker or causal selection process.

Research lineage

Decision 6 moves the citation diagnostic to passage meaning

A source can exist, be accessible and be cited while the selected passage remains poorly matched to the buyer's question. Semantic relevance is therefore a separate diagnostic stage.

Interpretation boundary

Three claims must remain separate

Final-output alignment
Supported for the selected Gemini fragment subset. Selected citation text is more aligned with its own query than with matched alternatives.
Internal relevance-gate validity
Not directly observed. Rejected passages, internal scores, thresholds and pass-or-fail decisions are unavailable.
Causal selection
Not established. The study cannot show that semantic relevance caused selection rather than authority, source quality, duplication control, policy, answer readiness or another downstream factor.

Observable passage proxy

Passage-like evidence exists only for a selected Gemini subset

Most final citations expose only a page or URL. Some Gemini citation URLs in the observed export contain browser text-fragment directives that point to selected text on the cited page. These anchors are more passage-like than ordinary URLs, but they do not establish the full passage considered internally, the text placed in model context or the candidate text evaluated by a reranker.

Bar chart showing passage-like text-fragment pointers in roughly 35% of cited Gemini answers, with the same citation format not observed in the corresponding ChatGPT and Perplexity exports.
Figure 3. Passage-like text-fragment pointers were observable in roughly one-third of cited Gemini answers and were not observed in the corresponding ChatGPT or Perplexity exports. Availability varied substantially over the historical collection, so the fragment subset is not representative of all citations.
Open full-resolution figure
Passage-like text-fragment evidence in the observed export
Platform Passage-like text-fragment evidence
ChatGPT Not observed in this export
Gemini ~35%
Perplexity Not observed in this export

Method

Matched alternatives create a demanding benchmark

A raw similarity score has no universal meaning across text representations or corpora. Decision 6 therefore uses a relative matched-decoy test focused on final-output query specificity.

  1. Pair the survivor

    Each selected fragment is paired with the originating prompt that produced the final citation.

  2. Choose five alternatives

    The fragment is compared with five fragments associated with different prompt templates.

  3. Keep decoys plausible

    Matching prefers comparable buyer-question and collection context, creating a harder negative control than unrelated text.

  4. Test specificity

    The analysis asks whether selected final-output text matches its own query more strongly than plausible wrong prompts.

The decoys are analyst-created controls. They were not observed as Gemini candidates and cannot be described as passages Gemini rejected.

Main quantitative result

Two text representations point in the same direction

The primary analysis uses an analyst-created latent-semantic text representation that captures broader word and phrase co-occurrence. Own-prompt similarity is approximately 0.27, versus approximately 0.09 for matched decoys—an advantage of about 0.18.

A secondary analyst-created lexical word-and-phrase representation produces a smaller alignment advantage of about 0.04. The selected fragment ranks first against five matched alternatives about 54% of the time under the semantic representation and about 51% under the lexical representation, compared with roughly 17% expected by chance.

Rounded matched-decoy alignment results
Representation Approx. alignment advantage Approx. top-1 rate
Latent semantic ~0.18 ~54%
Lexical ~0.04 ~51%
Chance benchmark ~17%

Interpretation: These are analyst-created relevance proxies applied to selected final citation fragments, not internal Gemini ranking scores. The meaningful result is comparative rather than a universal relevance threshold.

Buyer-question consistency

All five well-covered query categories show positive alignment

The aggregate result is not driven by one major buyer-question category. Every well-covered category shows positive own-prompt versus matched-decoy latent-semantic alignment.

Horizontal chart showing positive semantic-alignment advantages across five buyer-question types: about 0.23 for Product capability, 0.22 for Integration and ecosystem, 0.16 for Security and governance, 0.15 for Vendor evaluation, and 0.12 for Operational workflow.
Figure 2. All five well-covered buyer-question categories show positive own-prompt versus matched-decoy semantic alignment. Values are rounded public-safe estimates and should not be interpreted as internal platform relevance scores.
Open full-resolution figure
Rounded semantic alignment by well-covered buyer-question type
Buyer-question type Approx. semantic alignment advantage
Product capability ~0.23
Integration / ecosystem ~0.22
Security / governance ~0.16
Vendor evaluation ~0.15
Operational workflow ~0.12

The important result is that all five well-covered categories are positive, not that their exact ordering represents a durable ranking. Narrower query types do not have enough fragment-covered evidence for strong route-specific conclusions.

Robustness

Longer observable fragment text does not reverse the result

The positive semantic separation remains as the minimum amount of decoded fragment text increases and generally becomes stronger. Longer anchors give the analyst-created representation more evidence about what the cited text discusses.

This is a robustness check, not evidence of a hidden platform threshold. Fragment length is itself selected and may vary with query type, source type, answer format, interface behaviour or other unobserved factors.

Measurement boundary

Final-output relevance does not prove the hidden relevance gate

Highly relevant selected fragments tell us something important about the final output. They do not reveal how those fragments arrived there. Final citations may reflect semantic relevance, source authority, quality, freshness, duplication control, policy, formatting, answer usefulness and presentation logic.

Decision 6 does not measure Gemini's relevance model, does not show that Gemini uses either analyst-created representation and does not establish a reranker or threshold. The matched decoys are not observed candidates and were not observed being rejected by Gemini.

Missing denominator

We observe selected fragments, not the candidate population

The study begins at the final emitted citation. Rejected candidate passages are missing. This is survivor bias: we see passages that made it through the system, but not the population from which they were selected.

Without rejected candidates, the analysis cannot estimate sensitivity, specificity, false-positive rates, false-negative rates, threshold calibration, candidate-survival odds or the causal effect of semantic relevance on selection.

Supported conclusions

What Decision 6 establishes

  1. A testable passage proxy

    Gemini text-fragment pointers expose a selected-text proxy for a subset of final citations.

  2. Material own-prompt alignment

    Under the primary semantic representation, own-prompt similarity is roughly three times the matched-decoy level.

  3. Identification above chance

    Selected fragments rank first against five alternatives about 54% of the time, versus approximately 17% by chance.

  4. Independent directional support

    A simpler lexical analysis confirms the direction of the main result.

  5. Breadth across major categories

    All five well-covered buyer-question categories show positive semantic alignment.

Claim boundary

What Decision 6 does not establish

  • Internal ranking mechanics

    The study does not observe Gemini's internal relevance score, candidate set, reranker, threshold or rejected passages.

  • Causal selection

    It does not show that semantic relevance caused a citation to survive rather than another source or answer-quality factor.

  • Complete retrieved passage

    A decoded text fragment may not equal the complete passage retrieved or placed in model context.

  • Cross-platform generalisation

    The result does not apply to ChatGPT, Perplexity or Gemini citations without observable text-fragment evidence.

  • Evidence quality

    Passage relevance does not prove authority, quality, completeness or answer readiness.

Required telemetry

A direct test requires the candidate-level denominator

A direct candidate-level study would need every passage considered for each retrieval opportunity, along with its source identity, pre- and post-rerank position, retrieval and relevance scores, threshold status, survival decision, rejection reason, blinded human relevance label and final-answer outcome.

Candidate population
Every passage considered within a retrieval run, including the rejected passages missing from final-output data.
Ranking telemetry
Retriever and reranker scores, order changes, thresholds and structured pass-or-fail outcomes.
Independent labels
Blinded human relevance judgements that can be compared with platform scores and survival decisions.
Downstream outcome
Whether the candidate later appeared in the final answer and supported a useful response.

Stronger experiment

A confirmatory study should preserve complete candidate sets

  1. Sample distinct questions

    Cover the major buyer-question categories without treating repeated answers as independent evidence.

  2. Record every candidate

    Preserve passages before semantic filtering, including those that do not survive.

  3. Add blinded relevance labels

    Compare human judgements with retrieval scores, reranking and passage survival.

  4. Separate later filters

    Distinguish semantic relevance from source quality and answer readiness.

  5. Use interventions for causality

    Controlled changes are required to establish whether semantic match itself causes candidate survival.

AI visibility, AEO and GEO

A cited page should contain query-specific evidence

Citation presence is not enough. Teams should test whether important pages contain explicit passages that directly answer target buyer questions, use language aligned with how buyers ask them and retain their meaning when separated from surrounding marketing copy.

Important concepts should not be buried under unrelated material. Query-specific proof should be clear enough to stand on its own. These practices improve passage usefulness, but they do not guarantee citation, ranking or recommendation in a third-party AI system.

Semantic alignment is one layer of answer readiness. It is not a substitute for authority, factual quality, completeness or persuasive strength.

Relationship to Decision 5

Final-output properties remain downstream evidence

Decision 5 establishes that final citation output cannot prove historical page access or extraction. Decision 6 goes one step further: where passage-like final-output text is observable, it tests whether that selected text aligns with the query.

The same observability principle remains. A property of the final selected output is not automatically evidence of the hidden rule that selected it. Decision 5 applies that principle to extraction; Decision 6 applies it to semantic relevance.

Downstream diagnosis

Relevance is followed by quality and answer readiness

Once a passage is semantically relevant, later stages must ask whether the source is sufficiently authoritative, current, specific and trustworthy, and whether the surviving evidence can support a complete and useful answer.

Decision 7 asks whether a semantically aligned source and passage also show evidence consistent with surviving a later finalist-quality stage.

A passage may be relevant but weak. A high-quality source may not answer the exact question. Relevant, authoritative evidence may still omit information needed for a final answer. The later answer-readiness stage remains conceptual and is not linked until its production publication exists.

Study boundaries

Limitations

  • Selected Gemini subset

    Passage-like evidence is unavailable for most observed citations and was not observed in the corresponding ChatGPT or Perplexity exports.

  • Changing availability

    Text-fragment availability changed sharply across the historical collection, so the subset is not representative of all Gemini citations.

  • No rejected candidates

    Internal retrieval scores, reranker scores, thresholds and rejected passages are absent.

  • Imperfect passage proxy

    Decoded URL fragments may not equal the complete passage used internally or all surrounding context that affected selection.

  • Analyst-created controls

    The matched decoys and both text representations are public analytical proxies, not actual platform candidates or internal scoring systems.

Method and governance

Public-safe research details

This publication preserves the substantive matched-decoy finding while reporting rounded public-safe values. It describes roughly 6,000 usable Gemini prompt-fragment pairs across around 300 prompt templates without exposing reconstructive platform counts, fragment totals, route samples, collection trajectories, intervals or significance values.

The primary analysis uses an analyst-created latent-semantic representation. The secondary analysis uses an analyst-created lexical word-and-phrase representation. Neither should be attributed to Gemini, and neither reveals proprietary candidate-ranking logic.

Internal field names, file paths, analysis scripts, implementation details, collection-provider identity and customer-specific details are omitted. Structured data follows the same disclosure boundary as the visible article. No public candidate-level dataset accompanies the publication.

FAQ

Frequently asked questions

Do final AI citation passages match the query?

For the observable Gemini text-fragment subset, yes. Selected citation fragments are materially more aligned with their own prompts than with plausible matched alternatives, but this final-output result does not prove Gemini's hidden candidate-level relevance gate.

What is a passage-like text-fragment pointer?

It is a browser URL directive that points to selected text on a cited page. It provides a passage-like final-output anchor, but it does not prove what complete passage was retrieved, placed in model context or evaluated internally.

What did the matched-decoy analysis test?

It compared each selected Gemini fragment with its originating prompt and with five analyst-created fragments associated with different, contextually plausible prompt templates. The test asks whether selected final-output text is query-specific.

How much more aligned were selected fragments with their own prompts?

Under the analyst-created latent-semantic representation, own-prompt similarity was approximately 0.27 and matched-decoy similarity was approximately 0.09, an alignment advantage of about 0.18. The lexical alignment advantage was about 0.04.

How often did the selected fragment rank first against matched alternatives?

The selected fragment ranked first against five matched alternatives in roughly 54% of template-level comparisons under the latent-semantic representation and roughly 51% under the lexical representation, versus approximately 17% expected by chance. These are analyst-created comparisons, not Gemini candidate rankings.

Does this prove Gemini's internal relevance gate?

No. Rejected candidate passages, internal scores, thresholds and pass-or-fail decisions are unavailable, so the study cannot validate the hidden candidate-level relevance gate or show that semantic relevance caused final citation selection.

Why does the study use matched decoys?

Matched decoys provide plausible alternatives from comparable buyer-question and collection contexts. They create a harder benchmark than unrelated text and test whether the selected fragment retains query-specific meaning.

Does the result apply to ChatGPT and Perplexity?

No. The passage-like text-fragment convention was not observed in the corresponding ChatGPT or Perplexity exports, so the matched-decoy result is limited to the selected Gemini subset and does not describe how the other platforms perform passage selection.

Why is the Gemini fragment subset not representative of all citations?

Text-fragment availability changed sharply across the historical collection. The observable subset is selected by changing citation formatting, interface behaviour, collection conditions or related system behaviour and should not be treated as a random sample of all Gemini citations.

How does Decision 6 connect to Decisions 1–5?

Decision 1 studies retrieval need, Decision 2 studies buyer-question context, Decision 3 examines ecosystem source availability, Decision 4 examines selected-source composition, Decision 5 defines content-access and extraction observability, and Decision 6 tests passage-level semantic alignment in selected final citations.

Piush Vaish, founder and CEO of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish