Guide3 min readAI Design · Technology & Intelligence
Evals before prompts: how to test an AI feature before anyone relies on it
Without a fixed set of questions and agreed answers, every change to a prompt or a model is a matter of opinion. How we build eval sets, grade them, and turn them into a release gate.
Xterra Edze studioEditorial
Published

Teams building with language models often start with the prompt and judge it by trying a few questions. That works until a second person edits the prompt, the model is upgraded, or a new document enters the knowledge base, and nobody can say whether things got better or worse. An evaluation set, or evals, fixes that. It turns quality into a test that passes or fails.1 2
What goes in an eval set
- Real tasks. Questions and requests taken from the people who will use the system, not invented ones.
- The accepted answer. What a correct response contains, and the source it must cite.
- Refusals. Questions the system must decline or hand to a person: out of scope, unsafe, or not answerable from its sources.
- Attacks. Inputs that try to change the system’s instructions, including instructions hidden in documents it reads.3
- Edge cases. Long inputs, other languages, typos and ambiguous requests.
{
"id": "sso-pricing-07",
"slice": ["pricing", "high-risk"],
"input": "Does single sign-on cost extra on the Team plan?",
"must_cite": ["docs/plans#team"],
"accept": "Says whether SSO is included on Team, citing the plans page, and offers a person for pricing questions.",
"reject_if": ["quotes a price that is not in the sources", "gives no citation"]
}How to grade
Use the cheapest grader that is reliable for each check. Code-based checks cover a lot: did it cite a source from the allowed list, did it refuse, did it stay under the length limit. People grading against a rubric cover judgement. Model-based graders scale the rubric, but only after you have measured how often they agree with your human graders on the same cases.1

What a release report shows
- Pass rate per slice, against the target set before the run.
- New failures since the last release.
- Refusals that should have been answers, and the reverse.
- Injection attempts caught and missed.
- Cost and latency per answer.
Slices, not averages
A single score hides the failures that matter. Report results by slice: topic, risk level, language, customer segment. A release that improves general questions while getting worse on pricing or security answers is a regression, whatever the average says.
Make it a gate
Run the set automatically on every change: a new model version, a prompt edit, a new tool, new documents in the index. Set a pass threshold per slice before you see the results, and block the release when a slice falls below it. NIST’s AI Risk Management Framework puts this work under its Measure function: testing, evaluation, verification and validation, carried on after deployment rather than done once.4
Keep the set alive
- Add every real failure from production as a new case, with its correct answer.
- Retire cases that no longer reflect how people use the system.
- Re-check model-based graders against people on a regular schedule.
- Version the set with the code, so every result can be reproduced.
Open-source frameworks such as OpenAI’s Evals show the pattern in code.5 We design eval suites in AI Strategy & Agents and run them as release gates in AI Product & Automation.
Sources. Where the facts come from.
Numbered as they are cited in the text. Each link opens the original.
-
Define success criteria and build evaluations
Anthropic, Claude Platform Docs
Back to the text:Back to the text, place 1Back to the text, place 2





