Kojable research · Decision 8 · Answer readiness
Can AI-Cited Evidence Be Used Cleanly in the Final Answer?
Final cited answers commonly look direct, standalone-like and specific, but those downstream properties do not reveal whether the original finalist passages passed an answer-readiness gate or whether each citation supports the exact claim beside it.
Key finding
Across the three observed platforms, roughly nine in ten final cited answers satisfy the combined rendered-answer proxy, but this is a final-output property—not a passage-level Decision 8 pass rate.
Qualification: Finalist passages, rejected finalists, direct readiness labels, pass-or-fail outcomes and exact claim-to-evidence mappings are not observed.
- Answer readiness
- Claim-level support
- Citation placement
- 15 min read
Study overview
Executive summary
Decision 8 asks whether surviving evidence can support a clean, useful final answer. A source can be discoverable, accessible, relevant and strong enough to become a finalist while still requiring too much context, burying the answer or lacking the specificity needed for a precise claim.
Across more than 50,000 cited responses from ChatGPT, Gemini and Perplexity, final answers commonly begin with prompt-aligned, standalone-like and specific language under simple analyst-created proxies.
The approximate joint rendered-answer rates are 93% for ChatGPT, 93% for Gemini, 82% for Perplexity and 89% pooled. These are rendered final-answer surface properties, not Decision 8 passage pass rates.
Rendered cited answers show strong answer-readiness-like properties, but the finalist-level readiness gate and clean claim-level citation integration remain unproven because the historical export does not contain finalist failures, direct readiness labels or claim-to-evidence mappings.
Direct answer
Can AI-cited evidence be used cleanly in the final answer?
-
Strong final-output structure, unobserved passage gate
Final cited answers commonly exhibit prompt-aligned, standalone-like and specific openings, but those rendered properties do not prove that finalist passages passed an answer-readiness gate or that citations were mapped cleanly to the exact claims they support.
Final diagnostic stage
Decision 8 completes the evidence path
- Decisions 1–3
Retrieval need, buyer-question context and source availability establish whether external evidence is sought and present.
- Decisions 4–5
Selected-source composition and content-access observability examine which sources appear and what final URLs can reveal.
- Decisions 6–7
Passage semantic relevance and finalist quality move closer to the selected evidence without observing every hidden candidate decision.
- Decision 8
Answer readiness asks whether the surviving evidence can support a clean, specific and auditable final answer.
Final-stage concept
What answer readiness means
Answer readiness is distinct from relevance and finalist quality. A passage may be relevant and authoritative yet depend on several paragraphs of context, surface the answer indirectly or remain too general to support a precise statement.
Clean use also requires a claim-level relationship between the answer and its evidence. A source URL associated with an answer does not reveal which sentence, clause or factual assertion it supports.
Identification requirement
Direct proof requires finalist-level evidence
- Finalist denominator missing
The complete finalist passage population entering Decision 8 and the rejected finalists are unavailable.
- Criterion labels missing
There are no direct human labels for self-containedness, answer-first structure or specificity.
- Gate outcome missing
The Decision 8 pass-or-fail result is not present in the historical export.
- Claim mapping missing
Final-answer claim spans, citation offsets and claim-to-passage mappings are not observed.
Rendered-answer evidence ≠ passage-level gate evidence. A good final answer does not prove the original source passage was answer-ready.
Measurement boundary
Rendered answers are downstream transformations
A final answer is not a raw source passage. The model can paraphrase material, combine sources, reorder information, add connective language, compress long passages or move the conclusion to the beginning.
A concise, direct final sentence proves that the rendered answer has those properties. It does not prove the cited passage had them before answer generation.
Measurement framework
Three independent answer-readiness dimensions
- Self-contained
- Limited dependency on missing surrounding context, unresolved references or hidden setup.
- Answer-first
- The relevant point is surfaced early and with low reconstruction effort.
- Specific
- Concrete facts, entities, technical details or relationships support a precise claim.
These dimensions should remain separate. They do not form a public numerical Decision 8 score, and answer-first does not mean every page should use an FAQ format.
Transparent measurement
Final-answer proxies test observable surface properties
The analysis uses deliberately simple analyst-created rules to ask whether final-answer openings align with important prompt terms, read as relatively standalone and contain at least one concrete specificity signal.
The rules are not a trained quality model, an internal platform score or a human annotation system. Equal-template weighting keeps the descriptive comparison from being dominated by more frequent question templates.
Rendered-answer evidence
Most final cited answers satisfy the joint proxy
| Platform | Approx. rendered-answer joint proxy |
|---|---|
| ChatGPT | ~93% |
| Gemini | ~93% |
| Perplexity | ~82% |
| Pooled | ~89% |
These values measure rendered-answer surface properties under analyst-created rules. They are not internal platform scores or passage-level readiness pass rates.
Interpretation boundary
The roughly 89% proxy is not a gate pass rate
-
Final output, not finalist passage
The pooled proxy is measured after answer generation and includes only final survivors. It must not be described as “89% of passages pass Decision 8.”
ChatGPT and Gemini are approximately 93% under this rendered-answer proxy. That does not mean either platform has 93% answer-readiness accuracy. Perplexity is lower under this particular surface rule; the result does not establish lower underlying evidence quality.
Narrower source-side evidence
Selected Gemini anchors show partial passage-like signals
A selected Gemini subset includes browser text-fragment pointers that identify passage-like text. Within that subset, roughly three-quarters of observable anchors are standalone-like, roughly nine in ten contain a specificity signal and roughly four in ten satisfy the combined passage proxy package, including the already-public semantic-alignment criterion.
This is the closest available historical evidence to source-side answer readiness, but it remains selected and incomplete.
Scope boundary
The Gemini subset is not platform-wide evidence
- Historically selected
The fragment subset does not represent every Gemini citation.
- Platform-specific
The finding does not generalise to ChatGPT or Perplexity.
- Partial passage view
A browser fragment does not prove the complete internally used passage or surrounding context.
- Integration unavailable
The fragment does not reveal a Decision 8 outcome or exact use in the final answer.
Attribution boundary
Citation lists do not establish claim-level support
Source identity is observable; exact claim attribution is not. The approved export provides final source URLs separately from the rendered answer but no reliable answer-span offsets.
Why placement matters
Clean use requires evidence for the exact claim
- Partial sentence support
A citation may support only one clause while appearing beside a broader statement.
- Mixed claims
One citation may sit beside several unrelated assertions.
- Broad relevance
A source may concern the topic without supporting the exact wording.
- Overstatement
The rendered answer can express a stronger claim than the underlying evidence supports.
- Ambiguous attribution
Multiple listed sources may not map clearly to individual claims.
Exploratory comparison
Cited and no-citation answers do not identify the gate
Cited ChatGPT and Gemini answers tend to be longer and richer in technical or numeric signals than comparable no-citation answers. That broad relationship is descriptive and exploratory.
No-citation responses are not observed Decision 8 failures and may differ because of earlier stages in the retrieval and citation process.
Unit of analysis
Answer-first behaviour is difficult to infer upstream
A generated answer can move a conclusion to the beginning even when the source passage buries it. Observing answer-first structure in the final response therefore cannot prove the source passage was answer-first.
Decision 8 is conceptually about finalist passage usability. The historical proxy is measured on rendered response structure. Those units are not interchangeable.
Evidence supported
What Decision 8 establishes
- Rendered answers
- Most final cited answers exhibit prompt-aligned, standalone-like and specific openings under transparent proxies.
- Platform pattern
- ChatGPT and Gemini are near 93%, Perplexity near 82%, and the pooled rendered-answer proxy near 89%.
- Selected passage evidence
- Observable Gemini anchors provide partial, selected source-side readiness signals.
- Transformation boundary
- Final answer structure is not the same as source passage structure.
- Attribution boundary
- Citation URLs identify associated sources but not exact claim-level support.
- Gate status
- The direct finalist-level Decision 8 gate remains unproven.
Claims not supported
What Decision 8 does not establish
- No passage pass probability
Rendered-answer proxy rates do not measure the probability that a finalist passage passed Decision 8.
- No universal threshold
The proxies are transparent analyst rules rather than platform-internal scores or universal readiness standards.
- No inherited wording claim
A final answer's direct structure does not prove the source passage used the same structure.
- No complete attribution claim
A citation list does not demonstrate that each answer claim is supported cleanly.
Identification limit
Final-output modelling cannot recover the hidden gate
A concise, specific cited answer could come from a perfectly answer-ready passage, a long contextual passage rewritten by the model, several combined passages or general model knowledge paired with a broadly related source.
The same final output is compatible with several histories. A survivor-only model would learn what rendered cited answers tend to look like, not which passage properties caused a finalist to pass Decision 8.
Required evidence
Direct validation must record finalists and failures
- Retain the complete finalist set
Store accepted and rejected finalists with full passages and surrounding context.
- Label each readiness dimension
Measure self-containedness, answer-first structure and specificity independently.
- Record the Decision 8 outcome
Preserve pass-or-fail results and structured rejection reasons before answer generation.
- Map evidence to answer claims
Connect accepted passages to final claim spans using citation offsets and claim-to-passage mappings.
Future study
A stronger study would control passage properties directly
A confirmatory design would combine blinded human labels with complete finalist sets, record the gate outcome before answer generation, map accepted passages to final claims, evaluate calibration and error rates, and keep platform and model-version effects explicit.
Controlled passage edits could hold factual content constant while varying context dependence, answer position or specificity. That would provide stronger evidence about whether these properties influence finalist survival.
AEO and GEO implications
Make important evidence easy to use
- Stand alone
Write the important passage so it can be understood without reconstructing several sections of context.
- Surface the answer
Put the relevant conclusion early enough that the evidence has low reconstruction cost.
- Be concrete
Use precise entities, technical details, quantities or relationships where they support the claim.
- Support one clear claim
Keep evidence precise enough to quote or paraphrase without over-interpretation.
Immediate upstream stage
Decision 8 separates finalist quality from answer readiness
Decision 7 asks whether a semantically relevant candidate appears strong enough to become a finalist. Decision 8 asks whether that finalist can support the final answer cleanly.
High-quality evidence can still require too much context or fail to support one exact claim. The historical data show strong downstream answer properties but do not reveal the boundary between these hidden stages.
Complete series
The complete Decisions 1–8 diagnostic
- 1. Retrieval need
- 2. Buyer-question context
- 3. Source availability
- 4. Selected-source composition
- 5. Content access and extraction
Can source content be accessed and converted into usable text?
- 6. Passage semantic relevance
- 7. Finalist quality
Is the relevant source and passage strong enough to remain a finalist?
- 8. Answer readiness
Can the surviving evidence support a clean, specific final answer?
This is a diagnostic model for reasoning about where citation and representation failures can occur. It is not a claim that every AI platform uses these eight stages as proprietary architecture.
Series synthesis
A final citation cannot validate every upstream decision
-
The observability principle across Decisions 5–8
A final citation is the endpoint of several possible hidden decisions. It should not be treated as direct evidence for every upstream gate.
Source URL output is observable. Historical access, extraction, candidate relevance, finalist quality, passage readiness and claim-level integration often are not. Better measurement requires collecting the candidate denominator and the failures rather than inferring hidden mechanisms from downstream success.
Limitations
Interpret the downstream evidence narrowly
- Finalist population absent
There is no complete finalist set, rejected population or Decision 8 outcome.
- Generated answers
Final responses are downstream transformations rather than source-native passage text.
- Simple proxies
Analyst-created rules are transparent but are not direct human readiness labels.
- Selected source-side evidence
Passage-like evidence is limited to a selected Gemini subset and exact claim placement is unavailable.
Research governance
Public-safe research details
This publication preserves the substantive Decision 8 findings while removing reconstructive details from the historical collection. Public reporting uses rounded answer-level and passage-like proxy values.
Exact platform samples, citation-token totals, prompt-template and control counts, confidence intervals, effect sizes, sparse-subset results, collection dates, internal files, provider identity, customer identity and private review material are not published.
The diagrams communicate measurement boundaries and conceptual distinctions. They do not claim to reveal proprietary platform architecture.
FAQ
Frequently asked questions
What does AI answer readiness mean?
AI answer readiness asks whether surviving evidence can support a clean final answer. The public framework separates self-containedness, answer-first structure and specificity, while also requiring evidence to map clearly to the claims it supports.
Do final cited answers usually look answer-ready?
Yes, under transparent analyst-created proxies. Across more than 50,000 cited responses, the approximate joint rendered-answer rate is 93% for ChatGPT, 93% for Gemini, 82% for Perplexity and 89% pooled. These are final-answer surface properties.
Does the roughly 89% proxy rate mean 89% of passages pass Decision 8?
No. The proxy is measured after answer generation, contains only final survivors and can reflect rewriting, reordering, compression or combination by the model. It is not a passage-level Decision 8 pass rate.
What are the three answer-readiness dimensions?
Self-contained evidence makes sense with limited surrounding context, answer-first evidence surfaces the relevant point with low reconstruction effort, and specific evidence contains concrete facts, entities or relationships that support a precise claim.
Why is a rendered answer different from the source passage?
Answer generation can paraphrase, combine, reorder and compress source material or add connective language. A direct final answer therefore proves a property of the rendered response, not that the original source passage was answer-first or standalone.
Can citation URLs show which claim a source supports?
No. Citation URLs preserve source identity, but the historical export does not provide reliable answer-span offsets or claim-to-passage mappings showing which exact statement each source supports.
What does the Gemini passage-like subset show?
In a selected Gemini subset, roughly three-quarters of observable anchors are standalone-like, roughly nine in ten contain a specificity signal and roughly four in ten satisfy the combined passage proxy package. The selected fragments do not represent all Gemini citations or other platforms.
Why aren't no-citation answers valid Decision 8 failures?
A no-citation response may differ because retrieval was unnecessary or because discovery, access, relevance, finalist quality, policy or another earlier stage changed the outcome. No-citation controls are exploratory, not observed Decision 8 failures.
What evidence would directly validate an answer-readiness gate?
A direct study needs every finalist entering the stage, complete passages and context, separate readiness labels, accepted and rejected outcomes, final-answer claim spans, citation offsets and claim-to-passage mappings.
How does Decision 8 complete the Decisions 1–8 diagnostic?
Decision 8 closes a diagnostic sequence that moves from retrieval need and buyer-question context through source availability, source selection, content access, semantic relevance and finalist quality to answer readiness. It is a diagnostic model, not a claim about proprietary platform architecture.