A repeatable test set for answer presence, accuracy, citations, and drift

Rod Stockebrand
Co-founder, Brandleap.ai

Key Takeaways
Short on time? Here are the top things to know.
Article framework
What is an AEO evaluation dataset?
How many questions should a dataset contain?
What should each question record?
How do reviewers score generated answers?
How do you detect answer drift?
An answer evaluation dataset is not a pile of prompts someone copied into a spreadsheet. It is a measurement instrument. If the questions are biased toward easy branded lookups, the brand will look visible. If they omit comparisons, constraints, and customer language, the team will miss the places where buyers actually need evidence.
Treat the dataset like code: define a schema, review changes, preserve versions, and make failures reproducible. The objective is not to predict every possible prompt. It is to create a stable sample that represents the decisions your audience makes and the facts your organization must communicate correctly.
A good evaluation question has an expected answer, an owner who can validate it, and a reason it matters to the business.
Pull candidate questions from sales recordings, support tickets, site search, product forums, win-loss notes, and customer research. Preserve the customer’s language; normalize spelling only when needed for grouping. Tag every question by intent, funnel stage, product, segment, market, and importance. Include questions where the correct answer is conditional.
{
"id": "security-sso-07",
"question": "Which project tools support SAML SSO for a 100-person team?",
"intent": "implementation",
"importance": 5,
"market": "US",
"expectedFacts": ["SAML SSO", "100-person team", "project tools"],
"canonicalSources": ["/security/sso", "/pricing"],
"allowedQualifiers": ["plan-dependent"],
"reviewOwner": "Security product marketing"
}Expected facts are not a script the answer must repeat word for word. They define the propositions a correct answer should cover. Add acceptable qualifiers and prohibited claims when a fact is easy to overgeneralize. For pricing, require currency and plan scope. For compliance, require the relevant certification and its boundary rather than allowing “secure.”
Deduplicate near-identical prompts, but keep meaningful paraphrases when customers use different language. Balance branded and unbranded questions. Include your brand, competitors, category terms, problem statements, and “best” questions, while remembering that “best” requires a rubric. Sample across markets and languages if the business serves them; otherwise state that the dataset is intentionally narrow.
Duplicate versus useful variation
✗ Un-optimized
“Is Acme good?” / “Is Acme great?”
✓ Triple-rich rewrite
“Which tools support SAML for a 100-person team?” / “What should a 100-person team check before choosing SSO-enabled project software?”
Reviewers should not award one intuitive grade. Use labels that correspond to decisions: brand present or absent; expected fact correct, missing, or wrong; citation direct, partial, unrelated, or absent; answer complete, conditional, or overconfident; competitor present; sentiment or framing; and risk severity. A response can mention the brand accurately but cite a competitor for the decisive capability.
Write a short rubric with examples for every label. Run calibration: two reviewers score the same sample, discuss disagreements, and update the rubric before scaling. Track reviewer disagreement rather than hiding it. Ambiguity in the evaluation is a measurement finding.
Store the product, model or mode when exposed, collection timestamp, location, language, account state, prompt, complete answer, citations, and any tool or browsing indicator. Keep screenshots for interface evidence, but store structured text and URLs for analysis. Do not compare a logged-in personalized answer with an anonymous API response as if they were identical treatments.
Repeat a control panel more frequently than the full dataset. Answer products can vary because of retrieval, index updates, conversation state, and generation randomness. A single changed answer is an observation; a pattern across repeated controls is stronger evidence of drift.
When a question changes, create a new version and retain the old one. When a product feature or policy changes, update expected facts with an effective date. When reviewers change a scoring rule, do not silently rewrite historical labels. Report metric trends within the same dataset and explain breaks between versions.
dataset: brand-aeo-v3 rubric: answer-review-v2 collection: 2026-08-28T14:00Z question: security-sso-07 result: brand_present=true, accurate=partial, cited=direct, risk=low
For every failed question, link the error to an owner and a proposed intervention. A missing brand may require crawl or entity work. A wrong capability may require product documentation. An unsupported claim may require primary research. A stale answer may require change detection. Keep the original answer and rerun the question after the fix; otherwise the dataset becomes a retrospective rather than a feedback loop.
If every question comes from the loudest sales team or the most visible product, the dataset will overrepresent one part of the business. Create strata for intent, product, customer segment, geography, language, and risk. Set a minimum target for each important stratum, then reserve a small rotating sample for newly observed questions. This balances comparability with discovery.
Weighting can reflect revenue or risk, but show the unweighted counts too. A weighted score may be useful for prioritization, while an unweighted score shows breadth. Do not let a high-value segment disappear inside an aggregate, and do not let one large low-intent segment dominate the report simply because it supplied more prompts.
An answer does not need to repeat your approved sentence to be correct. Reviewers should compare propositions: the product supports the integration, the limit applies to the stated plan, the certification covers the relevant service, or the recommendation is conditional on the stated constraint. This prevents the dataset from rewarding memorization and makes it possible to recognize accurate paraphrases.
At the same time, define unacceptable shortcuts. “Enterprise-ready” is not an acceptable substitute for a required certification. “Real-time” is not acceptable when the documented behavior is hourly. A rubric should protect against confident generalities while allowing natural language variation.
Every expected fact should point to a source and a verification date. Every reviewer label should point to the answer span or citation that caused it. Every remediation should point back to the failed question. This creates a chain from customer need to public evidence to observed behavior. It also makes disagreements resolvable: the team can debate the source or the rubric rather than arguing from memory.
question -> expected fact -> canonical source -> observed answer
-> reviewer label -> owner -> content change -> retest
Some errors matter more than a missing marketing mention. Wrong legal, medical, financial, security, or eligibility information should receive a severity label and an escalation path. Add “appropriately uncertain” to the rubric. An answer that says it cannot verify a current policy may be preferable to a confident but outdated answer, especially when the canonical source is unavailable.
Include refusal and clarification cases when they are part of the customer experience. The evaluation question is not always whether the system answered; it may be whether it asked for the missing region, plan, version, or use case before recommending an option. Good evaluation measures usefulness and safety, not maximum answer length.
A question can become obsolete when a product is discontinued, terminology changes, or a market exits a region. Mark it retired with a reason rather than deleting it. Retired questions may remain valuable for historical analysis, but should not distort the current score. Add replacement questions when the underlying customer decision still exists under new language.
Review dataset health alongside answer metrics: coverage, duplicate rate, stale expected facts, reviewer agreement, collection failures, and time to remediate. A rising answer score is not reassuring if half the expected sources have not been checked in a year. The evaluation system needs maintenance just like the content it evaluates.
Primary references: NIST AI Risk Management Framework 1.0; Google Search Central guidance on AI features; and the RAGAS and HELM research communities’ public work on evaluating grounded and language-model systems. Product-specific answer behavior remains variable, so preserve raw observations.