Verification playbook
How to Verify Whether an AEO Change Worked
AEO verification is the post-intervention process of rerunning a predefined AI-answer measurement contract and comparing the results with a retained baseline to determine what moved, what held and what remains uncertain.
Here, AEO means Answer Engine Optimisation. The purpose is not simply to see whether an AI answer looks better after a change. It is to establish whether the specific representation gap you were trying to improve moved under comparable conditions, whether that movement is sufficiently consistent for the decision you need to make, and what the evidence actually permits you to conclude.
Quick answer
To verify whether an AEO change worked, rerun the exact frozen prompt panel against the same measurement contract used at T0, compare the predefined answer-alignment signals, and repeat observations when the conclusion requires evidence of stability. Keep platform-specific results separate and record material condition changes. Classify each target as Moved, Held, New Gap or Inconclusive.
A before-and-after difference shows observed movement. It does not by itself prove that the intervention caused the change.
If you do not already have a frozen T0 baseline and retest contract, start with the AI Visibility Tracking Baseline Guide before using this verification method.
Start here
Determine whether a predefined AI representation gap moved after a documented AEO or AI answer-alignment intervention, using the frozen T0 measurement contract.
- Goal
- Determine whether a predefined AI representation gap moved after a documented AEO or AI answer-alignment intervention, using the frozen T0 measurement contract.
- Inputs
- A retained T0 baseline, exact longitudinal prompt panel, defined AI surfaces and material test conditions, frozen run and measurement rules, intervention record, target representation gap and predefined movement rule.
- Output
- A retest observation record, signal-level before-and-after comparison, Moved / Held / New Gap / Inconclusive classification, evidence-strength statement and next operational action.
Comparison
Start with a valid comparison
Verification begins before you rerun a prompt.
You need to know what the intervention was intended to change and whether you retained enough pre-intervention evidence to make a meaningful comparison.
At minimum, confirm four things:
- A usable T0 exists.
You have the original prompt, AI surface, relevant test conditions and retained result.
- The target gap was defined before the retest.
You know what should appear, disappear or change if movement occurs.
- The intervention is documented.
You know what changed, where, when and what observable outcome it was intended to improve.
- The retest contract is known.
You know which prompts, surfaces, conditions, runs and outcome definitions are meant to be repeated.
A screenshot from three weeks ago and a new screenshot today can establish that two answers were different. It may not establish a defensible before-and-after measurement if the prompt, product surface, geography, mode or scoring rules are unknown.
Missing evidence does not always make comparison impossible. It limits the statement you can make.
For example:
If you retained the exact answer but not its citations, you may still assess an answer-level change but not a source-set change.
If the prompt text changed materially, you can compare the two answers descriptively, but you should not present the result as a clean longitudinal comparison of the same question.
If the original run design is unknown, a new multi-run retest cannot retrospectively establish the variability of T0.
The goal is to make the evidence carry no more certainty than the measurement design supports.
Prompt control
Keep the longitudinal panel frozen
The primary before-and-after panel should use the exact prompts retained at T0.
Do not quietly rewrite the question because a new wording seems more natural, more favourable or more likely to produce the desired answer.
Kojable's AI Visibility Tracking Baseline Guide treats exact prompt text as part of the measurement contract and versions meaningful wording changes rather than overwriting the original panel.
A 2026 paper accepted to ACM SIGIR found that generative-search source results were less consistent across repeated runs and less robust to minor query edits in the tested Google Search, AI Overview and Gemini environment.
This creates an important distinction.
Scroll horizontally if needed
| Test | Purpose | Prompt treatment | What the result tells you |
|---|---|---|---|
| Longitudinal retest | Did the original measured question change after the intervention? | Use the exact frozen prompt | Comparable before-and-after movement |
| Robustness test | Does the observed improvement extend beyond one wording? | Use separately predefined paraphrases or related prompts | Whether movement appears across wording variation |
| New buyer question | Has a new commercial information need emerged? | Add as a new prompt/panel version | A new measurement series, not historical T0 |
Robustness
Test paraphrase robustness separately
A paraphrase panel can be valuable. It simply answers a different question.
If the original prompt was:
Is [company] suitable for large enterprise teams?
a robustness panel might later include:
Is [company] a good fit for enterprise organisations?
That second prompt can test whether improved enterprise representation generalises across wording. It should not replace the original prompt in the historical comparison.
Keep the longitudinal and robustness panels separately labelled in the observation record. If a robustness prompt is introduced after T0, it has no historical T0 of its own unless an earlier comparable observation exists.
Timing
Set observation checkpoints, not propagation deadlines
There is no defensible universal rule that an AEO change should appear in AI answers after two weeks, four weeks or any other fixed interval.
Different interventions create different evidence-update paths. Updating an owned webpage, correcting a third-party directory, earning an independent review and publishing new research do not create the same discovery conditions.
Instead, define observation checkpoints in advance.
A checkpoint answers:
What is observable at this point?
It does not promise:
The change must have propagated by this point.
Google states that recrawling a changed URL can take from a few days to a few weeks and that requesting a crawl does not guarantee inclusion in search results. OpenAI separately documents OAI-SearchBot access as relevant to whether public content can be discovered for ChatGPT search. Neither source establishes a universal AI-answer update clock.
For example, a team might predeclare T+7, T+14 and T+30 as operating checkpoints. These are observation dates, not indexing, retrieval or answer-propagation guarantees.
When a result is unchanged at an early checkpoint, record what happened:
No qualifying movement was observed under the frozen retest conditions at T+7.
Do not automatically convert that observation into:
The intervention failed.
The evidence may justify a stronger conclusion later. An unchanged early result alone does not.
Conditions
Preserve the material test conditions
Repeating the same prompt is necessary for the longitudinal panel, but it may not be sufficient for a clean comparison.
Depending on the surface, material conditions can include:
provider;
product or AI surface;
selected model or mode where visible;
search or grounding state where controllable;
geography;
language;
fresh versus existing conversation;
account or personalisation conditions where material;
run design;
truth-set version;
metric definitions;
citation URL normalisation rules.
Kojable's AI Visibility Tracking Baseline Guide treats these conditions as part of the measurement environment.
You will not always be able to freeze everything. AI products change. Modes are retired. Model labels may disappear. Search behaviour evolves.
When that happens, version the measurement rather than hiding the change.
Scroll horizontally if needed
| Component | Longitudinal retest | If it changes |
|---|---|---|
| Exact prompt | Keep fixed | New panel/version |
| AI surface | Keep fixed where available | Record changed condition |
| Search/grounding state | Keep fixed where controllable | Record separately |
| Geography | Keep fixed where material | New subgroup or changed condition |
| Language | Keep fixed | New subgroup/version |
| Run design | Keep fixed | New measurement version |
| Metric definition | Keep fixed | Version the formula |
| Company truth set | Use the appropriate dated version | Record why company reality changed |
A changed environment does not make every comparison useless. It changes what the comparison means.
Evidence scope
Match providers and repetitions to the conclusion
There is no universal evidence-backed number of providers or runs that every AEO verification project must use.
The test design should follow the sentence you intend to say.
If your conclusion is:
One ChatGPT Search answer changed between T0 and this retest.
one valid observation can establish that dated observation.
If your conclusion is:
This representation change is recurring rather than an isolated output.
you need repeated comparable observations.
If your conclusion is:
The improvement appears across the AI surfaces our buyers use.
you need evidence from those relevant surfaces.
If your conclusion is:
AI systems now consistently represent the company differently.
the evidential burden becomes materially higher.
Kojable's Research, What Strong GEO Claims Require, reaches the same governing principle: stronger comparative and causal statements require stronger designs, including relevant opportunity, repetition, uncertainty and defined environments.
The IAB's August 2026 AI-visibility measurement framework similarly distinguishes directional measurement from measurement intended to support stronger decisions and highlights stability and reproducibility as core quality concerns.
The useful rule is:
Choose providers and repeated runs according to the scope of the intended conclusion.
Do not add platforms merely to reach an arbitrary minimum. Do not use one surface as evidence for a cross-platform claim.
Target signal
Compare the diagnosed gap, not a generic AEO score
A successful verification process begins with the gap that motivated the intervention.
Suppose AI answers repeatedly described a company as a mid-market tool even though its current positioning and evidence support enterprise use.
The primary verification question is not:
Did our AEO score increase?
It is:
Did the enterprise-positioning gap move?
That requires outcome measures tied to the diagnosis.
Kojable's AI Answer Accuracy and Alignment Guide separates factual accuracy, completeness, category and audience framing, competitive context and observable evidence rather than hiding them inside one unexplained score. The Research evidence parent similarly treats important outcomes as distinct rather than interchangeable.
Use the outcome that corresponds to the original problem.
Scroll horizontally if needed
| Diagnosed gap | Primary retest signal | Useful secondary evidence | What does not prove success by itself |
|---|---|---|---|
| Outdated company description | Outdated claim removed or corrected | Current supporting facts reflected | More citations |
| Wrong category | Correct category framing | Relevant category evidence present | Brand mention |
| Wrong audience | Current audience represented accurately | Relevant proof included | Longer answer |
| Missing capability | Required capability appears accurately | Supporting evidence/citation where relevant | Any new source |
| Weak enterprise proof | Required proof is represented | Evidence source appears where relevant | Citation count alone |
| Competitor displacement | Recommendation or comparison framing changes | Differentiation/proof reflected | Target company merely mentioned |
| Incorrect pricing/product fact | Incorrect fact disappears and current fact appears | Current source reflected | Improved sentiment |
| Source gap | Relevant source appears where source presence is itself the target | Source-set movement | Assumption that the new source caused the answer |
If you use a composite score, its formula, weights and decision purpose should be explicit. Otherwise the score can conceal the exact movement the verification exercise is supposed to detect.
Evidence
Treat citations and source changes as one evidence family
Citation movement can be useful.
You may observe that:
a company-owned page appears where it was previously absent;
an outdated third-party source disappears;
a new independent source appears;
source overlap between T0 and retest falls;
the provider cites a different set of URLs;
a source now supports a material claim that previously lacked visible evidence.
Those are legitimate observations.
They do not automatically establish:
that the source caused the answer;
that the platform considers it authoritative;
that citation frequency equals influence;
that the page was used in model training;
that a new citation means the overall answer is better aligned.
This matters because an answer can improve without citing your preferred source, and citation count can increase while factual representation gets worse.
Citation movement is therefore one evidence family inside answer alignment, not the verification result itself.
Decision
Classify the retest result
Every targeted gap should end in a clear operational state.
Scroll horizontally if needed
| State | Meaning | What you can say | Typical next action |
|---|---|---|---|
| Moved | The predefined target signal moved in the intended direction under the retest conditions | “The target signal moved under the stated conditions.” | Continue monitoring; consider whether broader validation is justified |
| Held | No qualifying movement was observed under the frozen conditions | “No qualifying movement was observed.” | Monitor again or return to diagnosis |
| New Gap | A materially relevant issue appears in the retest that was not part of the original target gap | “The retest exposed a separate representation issue.” | Diagnose the new gap |
| Inconclusive | Missing observations, changed conditions or variability prevent a defensible Moved/Held decision | “The available evidence does not support a clear movement decision.” | Repair or repeat the measurement |
These states deliberately describe what was observed, not why it happened.
Moved
A Moved result does not need to mean that every platform, run and answer became identical.
It means the predefined signal met the movement rule you established for that verification exercise.
If the target was removal of a specific outdated claim, for example, the claim might disappear consistently from the required retest cells while unrelated wording continues to vary.
Held
Held does not mean:
the source was unimportant;
the page was never discovered;
the intervention was ineffective;
the AI system ignored the evidence.
It means:
No qualifying movement was observed under the frozen retest conditions.
Why requires further evidence.
New Gap
Verification can expose an issue that was not previously salient.
For example, the enterprise description might improve while a newly surfaced answer introduces an incorrect implementation requirement.
Record that separately. Do not assume the original intervention caused it.
Inconclusive
Use Inconclusive when the evidence cannot support a clean movement decision.
Examples include:
too many planned observations failed;
the product surface materially changed between T0 and retest;
results vary too much to support the intended stability claim;
the original evidence was incomplete;
a scoring rule was changed after seeing the new answers.
“Inconclusive” is a measurement result. It is not an admission that the whole process failed.
Inference
Separate movement from attribution
This is the most important inference boundary in the Guide.
Suppose a company updates its enterprise product page on 1 August. On 30 August, repeated retests show that an outdated mid-market description is no longer appearing and enterprise evidence is now included.
You can directly say:
The tested answer pattern changed between the retained baseline and the August retest.
If that pattern recurs across the predefined observations, you may be able to say:
The movement recurred under the tested conditions.
You should not automatically say:
The product-page update caused the AI system to change its answer.
The intervention and the answer movement are temporally associated. Other parts of the public information environment may also have changed, as may the AI product itself.
Kojable's canonical Research on experimental design makes the distinction explicit. Descriptive observations require correct scope and denominators. Comparative claims need comparable opportunity, repetition and uncertainty. Causal claims require a design capable of separating cause from correlated alternatives.
A change log is still important. It establishes:
what changed;
when it changed;
where it changed;
who owned it;
what outcome was expected.
That creates temporal traceability.
It does not create causal identification by itself.
If several interventions happen simultaneously, the before-and-after design can still show whether the observed answer environment moved. Under this design, however, the individual intervention effects cannot be isolated without additional identification.
Illustration
Worked example: enterprise positioning
Consider an illustrative B2B software company moving from older mid-market positioning towards enterprise buyers.
This is an example of the method, not a customer result.
At T0, the company has already used the AI Visibility Tracking Baseline Guide to freeze four buyer-relevant questions:
- CAT-01
What are the leading [category] platforms for enterprise companies?
- FIT-01
Is [company] suitable for large enterprise teams?
- CMP-01
How does [company] compare with [competitor] for enterprise use?
- PRF-01
What evidence supports [company] for enterprise deployments?
The baseline records a recurring problem: several eligible answers continue to describe the company using its older mid-market framing, and relevant enterprise proof is absent.
The team diagnoses the gap and changes the evidence environment. For illustration, assume it updates the relevant enterprise page and adds verified deployment evidence.
Before making those changes, the retest contract has already been frozen.
At the first checkpoint, the team reruns the same prompt IDs on the same defined surfaces using the same run design.
The result record might look like this:
Scroll horizontally if needed
| Signal | T0 | Retest | Classification |
|---|---|---|---|
| Old mid-market framing | Present in qualifying observations | Still present | Held |
| Enterprise capability | Missing | Present in some observations | Moved for this signal |
| Enterprise proof | Missing | Results vary substantially | Inconclusive |
| Competitor recommended first | Present | Still present | Held |
| New factual issue | None | Incorrect implementation claim appears | New Gap |
This is more useful than declaring:
“The AEO change worked.”
Some dimensions moved. Others held. One is inconclusive. A new problem appeared.
The next action is therefore clearer:
preserve the enterprise capability improvement;
continue observing the proof signal;
diagnose why the competitor framing held;
investigate the new implementation error;
avoid claiming that the page update caused every observed change.
That is verification as an operating process rather than a pass/fail screenshot exercise.
Build a verification record
A good retest should leave enough evidence for another person to understand how the conclusion was reached.
For every observation, retain where relevant:
Verification record
prompt ID;
exact prompt;
prompt-panel version;
provider;
product/surface;
model or mode where visible;
search/grounding state where material;
date and time;
geography;
language;
run number;
valid/failed observation status;
full answer or retained answer record;
target company mentioned?;
targeted claim present?;
outdated claim present?;
category/audience aligned?;
relevant capability/proof present?;
competitors mentioned;
recommendation status;
visible citations;
source URLs/domains;
material changed conditions;
result state;
evidence-strength note;
next action.
The intervention record should sit next to the retest record:
what changed;
where;
date/time;
owner;
reason;
intended observable outcome;
dependencies or simultaneous changes.
Do not alter these definitions after seeing a favourable result unless the methodology change is recorded as a new version.
Verification checklist
Before the intervention
T0 baseline retained
Exact prompt panel frozen
AI surfaces defined
Material test conditions recorded
Run design frozen
Target representation gap defined
Success/movement rule defined
Metric or qualitative coding rules defined
Retest checkpoints set
Intervention record prepared
At each retest
Exact longitudinal prompts rerun
Same surfaces used where available
Run design preserved
Changed conditions recorded
Failed observations retained and classified
Full answers captured where permitted
Target signals assessed
Citation/source observations retained where relevant
Platform results reported separately before any aggregation
Robustness prompts kept separate from the longitudinal panel
Before reporting the result
Each target classified as Moved, Held, New Gap or Inconclusive
One favourable run has not been selected over conflicting evidence
Answer accuracy and citations have not been collapsed into the same claim
Changed measurement conditions are disclosed
The evidence strength matches the wording used
Temporal association is not described as causality without an appropriate design
The next Monitor, Diagnose or Improve action is recorded
Sources and further reading
Frequently asked questions
How do you know if AEO is working?
Start with the specific representation gap the AEO work was intended to improve. Rerun the frozen baseline panel and compare that target signal under the predefined retest contract. A meaningful result is not simply “more visibility”. It might be removal of an outdated claim, better category framing, inclusion of missing proof, improved recommendation context or a relevant source change.
Classify the result as Moved, Held, New Gap or Inconclusive, then state only what the evidence supports.
Should you use the same prompts when retesting AI answers?
Yes for the primary longitudinal panel. The exact frozen prompt gives you the cleanest comparison with T0.
Use paraphrases separately when you want to test whether the observed movement generalises beyond the original wording. Do not replace the longitudinal prompt with the paraphrase and present it as the same series.
How many times should you repeat an AI prompt when verifying a change?
There is no universally correct number of runs.
One run can document one dated observation. If you want to make a statement about repeatability or stability, you need repeated comparable observations. Higher-stakes decisions require stronger evidence and clearer treatment of uncertainty.
Preserve the run design between T0 and the retest.
Do you need to test multiple AI platforms?
Only when the scope of the intended conclusion requires it.
If the decision concerns one specific AI surface, that surface may be the correct analytical scope. If you want to claim that representation improved across several AI environments used by buyers, you need evidence from those relevant environments.
Do not assume one provider represents all providers, but do not add platforms merely to reach an arbitrary minimum.
How soon should you retest after an AEO change?
Set observation checkpoints in advance according to the decision, surface and expected evidence-update path.
Do not treat a checkpoint as a guaranteed propagation deadline. Google, for example, says recrawling can take days to weeks and does not guarantee inclusion after a crawl request.
An early unchanged observation should be reported as an observation, not automatically as failure.
Does a new AI citation mean the change worked?
Not necessarily.
A new citation is evidence that a visible source relationship changed in that observation. It does not by itself establish improved factual accuracy, better positioning, source authority, causal influence or successful answer alignment.
Compare the citation result with the representation outcome the intervention was intended to improve.
What does a Held result mean?
Held means:
No qualifying movement was observed under the frozen retest conditions.
It does not explain why.
The appropriate next action may be another checkpoint, a deeper diagnosis, or a different improvement path depending on the evidence and decision.
When should an AEO retest be classified as Inconclusive?
Use Inconclusive when the available evidence cannot support a defensible Moved or Held decision.
Common reasons include material changes in the test environment, insufficient repeated observations for the intended claim, missing T0 evidence, conflicting results or methodology changes that make comparison unreliable.
The next step is usually to repair or repeat the measurement rather than force a conclusion.
Can before-and-after AI answers prove that an intervention caused the change?
Not by themselves.
Before-and-after evidence can establish that an observable answer changed after an intervention. Repeated comparable observations can strengthen the evidence that the movement is recurring.
Causal attribution requires a design capable of isolating the intervention effect from plausible alternatives. See What Strong GEO Claims Require.
From verification to the next action
Verification is not the end of the process.
It determines what happens next.
Moved: retain the improvement, continue monitoring and decide whether broader validation is justified.
Held: continue observing if timing or variability remains relevant, or return to diagnosis if the evidence is strong enough to justify another action.
New Gap: record it as a separate issue and diagnose it rather than allowing it to disappear inside the original experiment.
Inconclusive: repair the comparison, collect the missing observations or create a new measurement version.
That returns the team to the wider operating loop:
Monitor → Diagnose → Improve → Verify → Monitor again.
Kojable uses this loop to help B2B companies understand and improve how AI represents them. Verification matters because improvement work should end with evidence about what changed, not an assumption that publication itself was success.
For the strategic context around the full operating process, see AEO Strategy. For ongoing representation monitoring, see AI Brand Monitoring. The Research evidence parent for the inference rules used in this Guide is What Strong GEO Claims Require.