AEOsim Blog
AI Answer Alignment Beyond Factual Accuracy: Six Observation Dimensions, Materiality and Recurrence Gates, and Misalignment Rates Across Prompt Families
Article 3/n on synthetic data in LLM visibility tools. See Article 1: How we Build the Prompt Set and Article 2: Bridge The Real2Sim Gap, not Sim2Real.
Table of Contents
Article 1 covered which questions to ask. Article 2 covered why a designed prompt set beats a harvested one. Neither tells you whether the answers coming back are any good. This article covers that.
Accuracy checks claims. Alignment checks the representation
An AI answer can pass every accuracy check and misrepresent the company anyway. Accuracy checks whether individual claims are correct. It says nothing about whether the overall picture (category, audience, competitive framing, evidence) matches current reality.
Accuracy asks: is this claim correct and current?
Alignment asks: does the overall representation match verified company reality and the available evidence?
A response can be accurate at the sentence level and misleading at the decision level. This happens when an answer places a company in an outdated category, targets the wrong segment, omits a capability the buyer's question depended on, compares it to alternatives on outdated criteria, or recommends a competitor for reasons the company's own current evidence contradicts.
Each sentence checks out. The combined picture steers the buyer wrong.
Acme Systems: accurate sentences, Diagnose disposition
Consider a fictional enterprise workflow platform, Acme Systems. Verified position: enterprise operations teams as primary audience, enterprise workflow platform as category, documented governance capabilities.
A buyer asks: "Which workflow platforms are suitable for a large enterprise that needs strong governance controls?"
The answer: "Acme Systems provides workflow automation and is commonly used by smaller and mid-market teams. Larger enterprises may prefer alternatives with stronger governance capabilities."
Nothing here is fabricated. Acme does provide workflow automation. Run it through an assessment and the cracks show.
| Dimension | Assessment |
|---|---|
| Factual accuracy | Product description accurate |
| Currency | Audience description outdated |
| Relevant completeness | Materially incomplete, current governance capabilities omitted |
| Category and audience framing | Misframed toward smaller and mid-market teams |
| Competitive framing | Distorted, alternatives favoured on a criterion Acme currently meets |
| Evidence support | Unavailable, no citation shown beside the claim |
| Materiality | Material, the buyer explicitly asked about enterprise governance fit |
| Recurrence | Observed once |
| Disposition | Diagnose |
A single accuracy score would land near "mostly correct." The assessment flags Diagnose despite recurrence reading "observed once," because this is a recommendation built on a limitation Acme does not have. One instance is enough to establish that kind of contradiction.
Six answer observations plus two action gates
Eight dimensions, not eight peers.
Six describe what was observed in one answer:
- Factual accuracy. Does the claim match current verified evidence? (Accurate / Inaccurate / Unverifiable)
- Currency. Is the claim still true as of the assessment date? (Current / Outdated / Uncertain / N/A)
- Relevant completeness. Does the answer include what matters for this buyer question, not the entire catalogue? (Sufficient / Materially incomplete / Uncertain)
- Category and audience framing. Is the company placed in the category and customer context it occupies today? (Aligned / Partial / Misframed / Unclear)
- Competitive framing. Is any comparison built on criteria that reflect current reality? (Aligned / Distorted / Unclear / N/A)
- Evidence support. Does the cited source actually support the claim beside it? (Direct / Partial / Unclear / Contradictory / Unavailable)
These six are assignable from the retained answer, its sources, and current evidence.
Two are different in kind:
- Materiality. Could this gap change the buyer's understanding or decision? (Material / Minor / Not material / Uncertain)
- Recurrence. Has the same gap appeared under comparable observations? (Observed once / Recurring / Unresolved)
Materiality depends on what the prompt was asking. That is a property of the question. Recurrence only exists across a set of observations. Neither can be read off one answer. They are gates: they decide whether a flagged observation is worth acting on.
Record the six per run. Compute the gates over the corpus. If all eight are recorded per answer, recurrence can only ever say "observed once."
Materiality should be fixed to the prompt's intent class before any answers are seen, not judged case by case. A completeness gap on a prompt with a hard constraint is material because the buyer named the constraint. The same gap on a definitional prompt is not. Deciding this in advance removes the assessor's discretion at the point where their incentive is worst.
This is also why a composite score does not work. A defensible one needs predefined weights, missing-data rules, and validation. Without those, one number hides whether the problem is a factual error, an omission, a framing issue, or weak evidence, and each needs a different fix.
Reviewer agreement and model judges per dimension
Article 1 required two independent reviewers on a sample, with recorded agreement rates. That applies here with more force, because these six judgments are softer.
They are not equally reliable. Factual accuracy, currency, and evidence support are close to mechanical: given a dated claim record and a cited page, two reviewers should converge. Completeness, category framing, and competitive framing require a view of what the buyer needed, and that is where reviewers diverge, and where an interested party's thumb lands on the scale.
Report agreement per dimension, not overall. High accuracy agreement next to near-chance completeness agreement tells you half the dimensions are opinion.
At scale, a model will be assigning these states. That is fine, but it is still a measurement with its own error: hold out a human-labelled subset, report agreement per dimension against it, and re-measure when the judge model changes. A judge upgrade can move your reported alignment rate without anything changing in how you are actually represented.
The Oumi study below graded AI Overviews using its own verification model, and Google's public objection was that one model was grading another on a benchmark with known errors. That objection stays fair, and the study's findings stay useful. Both hold when the model grading you is your own.
Correctness vs groundedness in cited answers
A citation is evidence to inspect, not proof the answer is accurate, and not proof the source caused the answer.
OpenAI's SimpleQA restricts itself to single-claim questions because correctness gets harder to judge as claims stack up. AttributionBench, an academic benchmark for claim-to-source verification, found that confirming whether a claim is fully supported by its citation remains unsolved.
Oumi's 2026 analysis of Google AI Overviews, run for the New York Times across 4,326 SimpleQA queries, found accuracy at 85% on Gemini 2 and 91% on Gemini 3. But 56% of the correct Gemini 3 answers were ungrounded, meaning the cited sources did not fully support the claim. Correct and supported together lands around 40%.
The direction matters more than the number. Ungrounded correct answers rose from 37% in October to 56% in February. Accuracy improved while verifiability got worse.
Correctness and source support are separate measurements. Treat them as separate steps.
Misalignment rate by family, engine, and cell
One retained answer is too thin a sample. Attach a rubric and it is still one draw.
Article 1 built prompt families for this reason: a canonical task with controlled variations, held out, reproducible from a recorded seed. Assessment inherits that structure. Run the family, not the prompt.
For each family, engine, and observation window, record the six observations per run and compute a misalignment rate per dimension:
rd = (runs where dimension d is flagged) / (runs in the cell)
Recurrence becomes a rate with a confidence interval instead of a guess. One screenshot gives r = 1.0 with an interval covering most of the range, the statistical form of the warning against treating one observation as a pattern.
The breakdown by cell is the diagnosis.
High on every engine. A category problem. The public corpus places you wrong, and no engine disagrees with all its sources at once. Content and positioning work, measured in months.
High on one engine, low on the rest. A source-mix problem specific to that engine's retrieval.
Middling on one engine with high variance. Not a position, an instability. The engine has not settled on a representation, and averaging across runs hides it completely.
The third pattern is the one a per-answer rubric cannot produce, and usually the one that changes what you do next.
Two carry-overs from Article 2: freeze an anchor block so rates stay comparable across a refresh, and keep families held out so a reported gap does not become a target.
Where to spend the monitoring budget
Article 2 argued a prompt is worth the uncertainty it reduces, peaking near even odds. Same shape here, with the outcome redefined as the per-dimension flag: information is proportional to r(1−r), maximised at r = 0.5.
A family misframed on every run, every engine, is a settled fact about the category. Worth establishing once, but it will not move week to week and will not tell you whether anything you shipped worked. A family never flagged is one you can stop running. The families near the middle are where a content or schema change can flip the outcome, and where you will actually see it flip.
One difference from Article 2: there, even odds meant your standing was genuinely uncertain. Here, a rate near the middle usually means the engine is unstable on this representation, which tends to be more actionable, since a stable misframe reflects the corpus and moves slowly, while an unstable one often reflects retrieval order and can move in weeks.
That gives a monitoring rule: run high-variance families often, run stable families on a fixed schedule, retire families that have not flagged anything across a full anchor period.
The same caveat from Article 2 applies. Monitoring only the informative middle biases your reported level. Keep a small unselected panel for the level, and use the selected set for movement.
Onset turn, correction persistence, elicitation survival
Article 2 relied on two things observation cannot do: forced elicitation and multi-turn control. A single-answer assessment has neither.
Buying conversations survey early and commit late. The turn where a buyer asks about governance controls is not the turn where the decision gets made. The Acme example is turn one of something longer. What does the answer look like at turn four, after the buyer has named a competitor or disclosed a constraint?
Trajectories add three things a single-answer rubric cannot record.
Onset turn. At which turn does the misframe first appear? Correct at turn one, dropped by turn four, is elimination under accumulating constraints. Misframed from the opening is category placement. Different problems.
Correction persistence. Introduce the correct fact as a user turn and keep going. Does it stick, or revert two turns later? A reverting answer is being rebuilt from the corpus each turn rather than carried in context, which means the fix is corpus-side, not conversational.
Elicitation survival. Article 2 found terminal turns demanding a single choice separate brands far more sharply than a mention grid. Run the same terminal turn at the end of a trajectory. Described accurately for four turns and excluded from the final recommendation is the highest-materiality failure this framework can find, and a single-turn check scores it Accept.
Trajectories also split three failures that look identical on one answer:
- Exclusion. The brand never enters the candidate set once constraints apply. Corpus and positioning work.
- Incoherence. The brand enters and leaves across repeated runs of the same trajectory. Grounding and source-consistency work.
- Extraction miss. The brand is in the retrieved sources with the capability stated, and the answer does not surface it. On-page structure and wording work.
Reporting "misaligned" without that split leaves you with a complaint instead of a diagnosis.
Hand-written multi-turn scripts do not scale, and they drift toward how a marketer imagines a buyer talks. Recent work on multi-turn user simulators tuned against preference signals (SMTPO, arXiv:2604.03671) is the relevant direction. Article 1's closed-loop warning still applies: if the simulator and the system under test share priors, you get the trajectories the model finds natural, not the ones buyers actually walk.
Accept, Monitor, or Diagnose
Every assessed answer ends in one of three states.
Accept. No material dimension flagged at rate, under the observed conditions.
Monitor. Material but below threshold, too few runs to estimate a rate, reviewers disagreed, or source support could not be assessed.
Diagnose. Flagged, material by intent class, and the rate clears a threshold over a minimum run count. Or, separately, a single observation directly contradicts a verifiable claim, which justifies review on its own regardless of rate.
Fix the threshold and minimum run count in advance, record them with the seed and weighting scheme, and do not move them after seeing results.
Five assessment mistakes that break the read
Treating visibility as accuracy. Appearing in an answer is not the same as being represented correctly within it.
Treating company preference as truth. Use current, supportable evidence, not whichever phrasing marketing would prefer.
Counting every omission as an error. Completeness is relative to the buyer's question, not an internal messaging document.
Treating citations as credibility scores. Inspect the relationship between claim and evidence directly.
Turning one screenshot into a pattern. A single observation can expose a real error. It cannot establish that the error recurs.
Limits: materiality, attribution, and gameability
Materiality is an assumption. The framework assumes a misframed answer changes what a buyer does. Proving that needs a counterfactual nobody has: what the buyer would have done given the aligned answer. Treat every materiality call as a plausibility argument.
There is no clean attribution. You cannot see the answer the buyer would have gotten without the gap, and once you have closed it you cannot cleanly separate your change from a model update or a competitor publishing the same week. Anchors narrow this. They do not close it.
The assessor is interested. Three of the six observations need a view of what the buyer needed, and the party running the assessment has a stake in that view.
The rubric is gameable. Content written to satisfy six dimensions will satisfy six dimensions, whether or not it helps the buyer. Held-out families are the only check, and only while they stay held out.
It does not compose into a score. Per-dimension rates show where representations break and how stable the breakage is. They do not add up to "well represented overall" without weights, missing-data rules, and validation nobody has built. Use the rates to find breakage and stability, not to invent an overall score.
Monitor, Diagnose, Improve, Verify
A material gap is only useful once verified: check the claim against current evidence, check whether it recurs, then diagnose, improve, and retest.
The sequence is Monitor, Diagnose, Improve, Verify, as separate stages. Verification is the stage most programmes skip, and the one that tells you whether the work did anything. Re-run the same family against the same engines with the anchor intact, and compare rates, not screenshots.
Frequently asked questions
What is the difference between AI answer accuracy and AI answer alignment?
Accuracy evaluates whether individual claims are correct and current. Alignment evaluates whether the overall representation, including completeness, framing, and competitive context, matches verified company reality for the specific question asked. Sentence accuracy does not guarantee overall alignment.
Does a missing fact automatically make an answer inaccurate?
No. An omission matters when it changes the buyer's understanding of fit, limitation, or differentiation for their specific question, not simply because the fact exists in an internal document.
Does a citation mean an answer is accurate?
No. A cited source can directly support a claim, partially support it, fail to support it, or contradict it outright. A source appearing beside an answer does not establish that it caused the answer.
Is one incorrect answer enough to justify action?
Depends on the kind of error. A single observation can establish a direct factual contradiction and justify immediate review. It cannot establish that the problem recurs, is provider-wide, or was caused by a specific source. Those need evidence across multiple runs.
How many runs before a misalignment rate means anything?
Enough that the confidence interval is narrower than the decision you are making, which depends on how unstable the engine is on that family. Report the interval alongside the rate and let the reader see how much room it leaves.
Can a model do the assessment instead of a person?
Yes, and at any real volume it has to. Treat its labels as measurements with error, not ground truth: hold out a human-labelled subset, report agreement per dimension against it, and re-measure when the judge model changes.
References
- OpenAI (2024). Measuring short-form factuality in large language models (SimpleQA).
- Li, Y., Yue, X., Liao, Z., & Sun, H. (2024). AttributionBench: How Hard is Automatic Attribution Evaluation?.
- Oumi (2026). Analysis of Google AI Overviews accuracy and groundedness, conducted for The New York Times.
- SMTPO: multi-turn user simulator preference optimisation. arXiv:2604.03671.
Vignesh Kanike is Founding Engineer at AEOsim, and Piush Vaish is founder of Kojable. This is Article 3/n on synthetic data in LLM visibility tools. Read Article 1 and Article 2. On how clarifying answers reshuffle the recommendation set, see Clarification as a Branch Point. Measure multi-turn AI visibility with Analytika.