Assessment playbook
How to Assess AI Answer Accuracy and Alignment
Assess AI answer accuracy by comparing the material claims in a captured response with a dated, verified company truth set. Accuracy is only one part of AI answer alignment, which asks whether the overall representation matches current company reality and relevant evidence. Check factual accuracy, relevant completeness, category and audience framing, competitive context and observable citation support. Then judge whether any gap is material and whether it recurs under comparable observations. The result should be a transparent decision such as Accept, Monitor or Diagnose, not an unexplained overall accuracy score.
This Guide starts after you have captured the answer. If you still need to define buyer questions, AI surfaces, run conditions and T0, use Kojable's AI Visibility Tracking Baseline Guide first. That Guide establishes the measurement layer before diagnosis or improvement.
Start here
Determine whether a captured AI answer is materially aligned with current verified company reality and evidence.
- Goal
- Determine whether a captured AI answer is materially aligned with current verified company reality and evidence.
- Inputs
- A captured AI answer, the exact buyer question, platform and product surface, observation date and run identifier, a versioned company truth set, and visible citations where available.
- Output
- A transparent alignment assessment record ending in Accept, Monitor or Diagnose.
AI answer accuracy is only one part of answer alignment
AI answer accuracy asks whether the factual claims in an AI response are correct and current.
AI answer alignment is broader. It asks how well the observable AI representation matches verified company reality, current positioning and relevant available evidence.
A response can therefore be factually accurate and still be materially misaligned.
For example, an AI answer might correctly describe what a software product does while:
placing it in an outdated category;
associating it with the wrong customer segment;
omitting a capability that is essential to the buyer's question;
comparing it using criteria that do not reflect its current market position;
recommending an alternative for reasons contradicted by current evidence.
Kojable's AI Representation reference entry defines representation as the observable pattern of how AI systems describe, categorise, compare, cite and recommend a company. It also distinguishes visibility from alignment: appearing frequently does not establish that the company is being represented accurately or usefully.
This Guide does not repeat that conceptual framework. It turns it into a practical assessment.
Before you assess an answer, establish the reference point
An AI answer cannot be assessed against a vague idea of what the company would prefer the model to say.
You need a verified company truth set: a dated, versioned reference containing the current facts, positioning, limitations and evidence relevant to the buyer question being assessed.
At minimum, retain:
the exact buyer question;
the captured answer;
the AI provider and product surface;
the observation date;
the run or observation identifier;
the truth-set version;
the company attributes relevant to that particular question;
visible citations or sources, where the surface provides them.
The truth set should distinguish current verified information from internal aspiration.
For example, "we want to be perceived as an enterprise platform" is not enough. The assessment should be grounded in evidence such as the company's actual customer scope, product capabilities, documentation, current positioning and supportable proof.
The existing baseline methodology treats the truth set as part of the measurement contract and requires it to be versioned when company reality changes.
The important practical rule is:
Do not compare every answer with every fact the company knows. Compare it with the facts and evidence materially relevant to the buyer question.
If the important buyer questions and their required evidence have not already been mapped, use the AEO Buyer-Question Mapping Playbook before deciding which omissions are material.
The AI answer alignment assessment framework
Use the following dimensions separately before making an overall operational decision.
Scroll horizontally if needed
| Dimension | Core question | Assessment states |
|---|---|---|
| Factual accuracy | Does the material claim match current verified evidence? | Accurate / Inaccurate / Unverifiable |
| Currency | Is the claim current for the assessment date? | Current / Outdated / Uncertain / N/A |
| Relevant completeness | Does the answer contain what materially matters for this buyer question? | Sufficient / Materially incomplete / Uncertain |
| Category and audience framing | Is the company placed in the appropriate category and customer context? | Aligned / Partial / Misframed / Unclear |
| Competitive or recommendation framing | Is the comparison or recommendation contextually appropriate? | Aligned / Distorted / Unclear / N/A |
| Evidence support | Does visible cited material support the relevant claim? | Direct / Partial / Unclear / Contradictory / Unavailable |
| Materiality | Could the gap materially alter the buyer's understanding or decision? | Material / Minor / Not material / Uncertain |
| Recurrence | Has the same defined gap appeared under comparable observations? | Observed once / Recurring / Unresolved |
| Disposition | What should happen next? | Accept / Monitor / Diagnose |
This is an assessment framework, not a validated numerical scoring scale.
Converting these dimensions into one number would require prospectively defined weights, missing-data rules, analytical units and validation. Without those controls, a single "alignment score" can hide important differences between factual error, omission, framing and evidence quality.
Claims
Identify the material claims in the answer
Start by separating the answer into the claims that matter to the buyer decision.
Do not score every sentence mechanically. Focus on statements that affect how the reader would understand:
what the company is;
who it serves;
what it can or cannot do;
where it fits;
how it differs from alternatives;
whether it meets the buyer's stated requirements;
what evidence supports an important conclusion.
A sentence can contain several claims.
For example:
"Acme is a mid-market compliance platform designed primarily for smaller financial teams."
contains at least three assessable propositions:
Acme is a compliance platform.
Acme is positioned for the mid-market.
Its primary audience is smaller financial teams.
One part can be accurate while another is outdated.
Breaking the answer down this way prevents a mostly correct paragraph from hiding one commercially important error.
The principle has precedent in factuality research. OpenAI's SimpleQA benchmark deliberately uses short questions with single, verifiable answers because factual correctness becomes harder to evaluate as responses contain more claims and ambiguity. The benchmark is not a B2B representation framework, but the methodological lesson is useful: define what is being judged before judging it.
Facts
Check factual accuracy and currency
Compare each material factual claim with the relevant evidence in the truth set.
Use a small set of explicit states.
Accurate
The claim agrees with current verified evidence.
Inaccurate
The claim directly conflicts with verified evidence.
Outdated
The claim was previously supportable but is no longer current.
Unverifiable
The available evidence is insufficient to classify it confidently.
Do not force uncertainty into "wrong".
An unverifiable statement may expose an evidence problem, but it is different from a statement that demonstrably contradicts the company's current reality.
Currency deserves separate attention because AI representation is dated. A description that was correct two years ago can be wrong today without ever having been fabricated.
This is why the assessment should record both the answer date and truth-set version.
Completeness
Judge relevant completeness, not total completeness
Completeness is one of the easiest assessment dimensions to misuse.
An AI answer does not need to repeat the entire product catalogue, company history or approved messaging framework to be aligned.
The correct question is:
Did the answer omit something materially necessary to answer this buyer question properly?
Suppose a company offers:
single sign-on;
audit logging;
regional data hosting;
workflow automation;
dozens of integrations;
several reporting features.
If the buyer asks:
"What does this company do?"
omitting regional data hosting may be perfectly reasonable.
If the buyer asks:
"Which platforms meet our enterprise security and governance requirements?"
the same omission could materially change the apparent fit.
Kojable's AI Representation framework makes the same distinction: capability completeness should be assessed against prompt intent, not against a requirement that every answer reproduce the whole product catalogue.
Use:
Sufficient
The answer includes the information necessary for the buyer decision.
Materially incomplete
A relevant omission changes or could reasonably change the understanding of fit, limitation, proof or differentiation.
Uncertain
It is not clear whether the omitted information was required for the question.
This is more defensible than counting missing brand messages.
Framing
Assess category, audience and competitive framing
Factual correctness alone does not tell you whether the answer creates an accurate overall picture.
Check how the individual claims combine.
Category framing
Does the answer place the company in the category it currently operates in?
A company may be named correctly while being repeatedly associated with an earlier category.
Audience framing
Does the answer represent who the company is actually designed to serve?
A specialist enterprise provider can be described accurately at product level while still being framed as a small-business tool.
Use-case framing
Does the answer connect the company to use cases it genuinely supports?
Likewise, does it imply suitability for a use case the company does not actually serve?
Competitive framing
When alternatives are named, examine the criteria used to compare them.
The assessment question is not whether the company "wins".
It is whether the comparison reflects current evidence and applies appropriate criteria to the buyer's question.
Recommendation fit
A recommendation is stronger than a mention.
If the AI system recommends or excludes the company, capture the reason given and compare that reasoning with current evidence.
The assessment can conclude:
The comparison materially misstates the company's current enterprise fit.
It should not jump to:
Competitor content caused the model to frame the company this way.
If the observed problem is specifically competitor-led framing and needs deeper investigation, continue to the Competitive Gap Audit Guide. That Guide owns competitor-specific diagnosis rather than this assessment step.
Evidence
Check what the citations actually support
Visible citations deserve systematic inspection, but citation presence is not the same thing as answer accuracy, evidentiary support or causal influence.
A useful sequence is:
Identify the material claim.
Identify the visible citation associated with it, where available.
Open the cited source.
Check whether the source directly supports the claim.
Record the relationship.
Use states such as:
Direct support
The cited evidence clearly supports the material claim.
Partial support
The evidence supports part of the claim, but not all of it.
Unclear
The connection is ambiguous.
Contradictory
The cited evidence conflicts with the answer's claim.
Unavailable / not assessable
The surface provides no usable source evidence, or the source cannot be assessed reliably.
Do not infer more than the observation allows.
Kojable's Citation Count Is Not Citation Quality Research separates citation volume, material-claim citation coverage and reviewed evidentiary support because these are different measurements. For the broader definition and measurement boundaries, see AI Citations.
Academic attribution research makes the same methodological distinction: AttributionBench shows that evaluating whether generated claims are fully supported by cited evidence remains a non-trivial problem.
A 2026 Oumi study of Google AI Overviews provides a useful illustration, although it should not be treated as a B2B benchmark. In its SimpleQA-based analysis, about 91% of assessed overviews contained the correct target answer, while 39% were both correct and fully supported by the displayed sources under Oumi's evaluation method. The study therefore demonstrates why correctness and source support should be assessed separately.
The practical rule is simple:
A citation is evidence to inspect, not proof that the answer is accurate and not proof that the source caused the answer.
Materiality
Decide whether the gap is material
Not every discrepancy deserves a project.
Materiality asks whether the observed difference could reasonably change the buyer's understanding or decision in the context of the question.
A gap is more likely to be material when it changes understanding of:
the company's category;
its primary audience;
suitability for the requested use case;
an important capability or limitation;
trust or proof relevant to the decision;
the criteria used in a competitive comparison;
the basis for inclusion or exclusion from a recommendation.
A minor wording difference may be immaterial even if the company would have written the sentence differently.
Likewise, absence of a preferred marketing phrase is not automatically a representation problem.
The purpose of this step is to prevent teams from turning every disagreement with an AI answer into an optimisation task.
A useful classification is:
Material
Minor
Not material
Uncertain
Where uncertainty remains, record it rather than manufacturing precision.
The IAB 2026 AI visibility measurement framework makes a related methodological distinction between directional signals and evidence that is sufficiently stable and reproducible to support stronger decisions. The same principle applies here: the strength of the action should match the strength and relevance of the evidence.
Recurrence
Separate one observation from a recurring pattern
A single answer can establish a direct factual contradiction.
For example:
The answer says the company no longer supports Product X. The current product documentation shows that Product X is supported.
That observation can be recorded immediately.
But one answer cannot by itself establish:
that the problem is persistent;
that it appears across buyer questions;
that it is stable on the same surface;
that another provider behaves the same way;
that one source caused it.
The existing Kojable baseline methodology treats one run as one observation and recommends designing repeated runs around the decision being measured rather than inventing a universal run count.
Classify recurrence separately:
Observed once
The defined gap appears in one retained observation.
Recurring
The same material gap appears under comparable observations.
Unresolved
The current evidence does not support either conclusion confidently.
Recurrence strengthens the case for further diagnosis, but it should not become a rule that a serious factual error must be ignored until it repeats.
A direct, high-impact false statement can justify immediate review while still being labelled as one observation.
Disposition
Classify the result as Accept, Monitor or Diagnose
The assessment should end with a decision, not another dashboard number.
Kojable uses Answer Intelligence as the diagnostic capability for investigating recurring answer patterns, source associations, outdated information, missing proof, competitor framing and likely drivers without presenting observational evidence as exact causality.
The operating sequence remains:
Monitor → Diagnose → Improve → Verify
Worked example: the facts are mostly right, but the answer is still misaligned
Consider a fictional B2B software provider, Acme Systems.
Its verified current position is:
primary audience: enterprise operations teams;
current category: enterprise workflow platform;
core product function: workflow automation;
enterprise deployment and governance capabilities are current and publicly documented.
A buyer asks:
"Which workflow platforms are suitable for a large enterprise that needs strong governance controls?"
The captured AI answer says:
"Acme Systems provides workflow automation and is commonly used by smaller and mid-market teams. Larger enterprises may prefer alternatives with stronger governance capabilities."
At first glance, the answer is not entirely wrong. Acme does provide workflow automation.
Apply the framework:
Scroll horizontally if needed
| Dimension | Assessment |
|---|---|
| Factual accuracy | Product description accurate |
| Currency | Audience description outdated |
| Relevant completeness | Materially incomplete because current governance capabilities are omitted |
| Category/audience framing | Misframed towards smaller and mid-market teams |
| Competitive framing | Potentially distorted because alternatives are favoured on a criterion Acme currently supports |
| Evidence support | Inspect visible citations claim by claim; citation presence alone does not resolve the discrepancy |
| Materiality | Material because the buyer explicitly asked about enterprise governance fit |
| Recurrence | Undetermined until comparable observations are checked |
| Disposition | Diagnose if the factual/current evidence is clear; recurrence should still be measured separately |
The key insight is that the answer is not simply "70% accurate" or "mostly correct".
Its basic product description is accurate, but the representation presented to this particular buyer is materially misaligned.
That is why one overall score can be less useful than explicit assessment states.
Common assessment mistakes
Treating visibility as accuracy
A company can appear in every relevant response and still be represented inaccurately.
Visibility answers:
Did we appear?
Alignment asks:
Was the representation sufficiently accurate and appropriate?
Treating company preference as truth
The assessment is not a mechanism for marking every sentence that differs from preferred marketing language as wrong.
Use current, supportable evidence.
Counting every omission as an error
Completeness is question-relative.
An omitted fact matters when it changes the answer to the buyer's question, not merely because it appears in the internal messaging document.
Treating citations as credibility scores
A cited answer is not automatically more accurate.
Inspect the relationship between the material claim and the evidence.
Turning citations into causal source attribution
A source appearing beside an answer does not establish that the source caused the answer, was the most influential input, was trusted by the model or was used in training.
Treating one screenshot as a stable pattern
One observation can expose an error. It cannot establish persistence.
Record both the direct observation and recurrence status.
Creating an overall score before defining the measurement
A composite alignment score needs explicit components, weights, missing-data rules, analytical units and validation.
If those do not exist, categorical assessment is more transparent.
A practical AI answer alignment record
For every assessed answer, retain:
Workbook output
Buyer question
Platform and surface
Date and run
Truth-set version
Material claims
Factual accuracy
Currency
Relevant completeness
Category/audience framing
Competitive/recommendation framing
Evidence support
Materiality
Recurrence
Disposition: Accept / Monitor / Diagnose
Reviewer notes and evidence links
This creates an inspectable record that another reviewer can challenge or reproduce.
It also makes later retesting more meaningful because the team knows exactly what gap it is checking.
Sources and further reading
Citation Count Is Not Citation Quality — Kojable Research on citation volume, material-claim coverage and reviewed source support.
AI Visibility Tracking Baseline Guide — prerequisite measurement contract for prompts, surfaces, runs and T0.
AI Representation — conceptual definition and high-level representation dimensions.
AI Citations — citation terminology and measurement boundaries.
OpenAI SimpleQA — factuality benchmark methodology.
AttributionBench — claim-to-cited-evidence attribution evaluation.
Oumi: Google AI Overview factuality and source support — current illustration of correctness and support as separate measurements.
IAB: Measuring Visibility in the AI Era — methodology disclosure, stability and reproducibility guidance.
Frequently asked questions
What is the difference between AI answer accuracy and AI answer alignment?
AI answer accuracy concerns whether factual claims in the answer are correct and current. AI answer alignment is broader. It assesses whether the overall representation matches verified company reality and relevant evidence, including completeness, category and audience framing, competitive context and other factors that matter to the buyer question.
A response can therefore be factually accurate in its individual statements while still creating a materially misleading picture of the company.
What is a verified company truth set?
A verified company truth set is a dated and versioned reference containing the current facts, positioning, scope, limitations and evidence used to assess an AI answer.
It should contain supportable information rather than simply the wording the company would prefer an AI system to repeat.
For each assessment, use only the attributes relevant to the buyer question.
Does every missing company fact make an AI answer inaccurate?
No.
An answer does not need to contain every known capability or proof point.
An omission becomes relevant when the missing information is materially necessary to answer the buyer's question or understand the company's fit, limitation, evidence or differentiation.
That is why this Guide uses relevant completeness, not total completeness.
Do citations mean an AI answer is accurate?
No.
Citation presence and factual accuracy are different measurements. A cited source may directly support a claim, partly support it, fail to support it or even contradict it.
Likewise, seeing a source attached to an answer does not establish that the source caused the answer.
Assess the claim and the evidence relationship separately.
Is one incorrect AI answer enough to take action?
It depends on the error and the intended action.
One retained response can establish a direct factual contradiction. A serious current error can justify immediate review.
However, one observation does not establish recurrence, provider-wide behaviour or causal explanation. If those conclusions matter, collect comparable evidence before making them.
Should AI answer alignment be expressed as one score?
Not automatically.
A defensible composite score requires defined components, weights, analytical units, missing-data rules and validation.
Without those controls, separate categorical judgements can show far more clearly whether the problem is factual accuracy, completeness, framing, evidence support, materiality or recurrence.
What should a team do after finding a material alignment gap?
Move from assessment into diagnosis.
Verify the evidence, examine whether the gap recurs, review the source and information patterns associated with it, then determine what action is justified and what should later be retested.
Kojable's role is to connect that process through Monitor → Diagnose → Improve → Verify, while keeping observation, likely influence and demonstrated change separate.
Move from a material gap to a justified action
Finding the gap is useful only if the next step matches the evidence.
Kojable helps B2B teams move from an observed representation problem into evidence-backed diagnosis, prioritised improvement guidance and comparable retesting.