Guide3 min readAI Design · Technology & Intelligence

Evals before prompts: how to test an AI feature before anyone relies on it

Without a fixed set of questions and agreed answers, every change to a prompt or a model is a matter of opinion. How we build eval sets, grade them, and turn them into a release gate.

Xterra Edze studioEditorial

Published

Close-up of the scale on a steel vernier caliper.
Quality you can measure, to the same scale every time.Photo: Bozhin Karaivanov on Unsplash

Teams building with language models often start with the prompt and judge it by trying a few questions. That works until a second person edits the prompt, the model is upgraded, or a new document enters the knowledge base, and nobody can say whether things got better or worse. An evaluation set, or evals, fixes that. It turns quality into a test that passes or fails.1 2

What goes in an eval set

  • Real tasks. Questions and requests taken from the people who will use the system, not invented ones.
  • The accepted answer. What a correct response contains, and the source it must cite.
  • Refusals. Questions the system must decline or hand to a person: out of scope, unsafe, or not answerable from its sources.
  • Attacks. Inputs that try to change the system’s instructions, including instructions hidden in documents it reads.3
  • Edge cases. Long inputs, other languages, typos and ambiguous requests.
json
{
  "id": "sso-pricing-07",
  "slice": ["pricing", "high-risk"],
  "input": "Does single sign-on cost extra on the Team plan?",
  "must_cite": ["docs/plans#team"],
  "accept": "Says whether SSO is included on Team, citing the plans page, and offers a person for pricing questions.",
  "reject_if": ["quotes a price that is not in the sources", "gives no citation"]
}
One case from an eval set. The rubric is written before the prompt, by someone who knows the answer.

How to grade

Use the cheapest grader that is reliable for each check. Code-based checks cover a lot: did it cite a source from the allowed list, did it refuse, did it stay under the length limit. People grading against a rubric cover judgement. Model-based graders scale the rubric, but only after you have measured how often they agree with your human graders on the same cases.1

A person typing on a laptop in a dark room, code on the screen, a lamp behind.
Evals run in the pipeline, not in someone’s chat window. Photo: Bryan Plata on Unsplash

What a release report shows

  • Pass rate per slice, against the target set before the run.
  • New failures since the last release.
  • Refusals that should have been answers, and the reverse.
  • Injection attempts caught and missed.
  • Cost and latency per answer.

Slices, not averages

A single score hides the failures that matter. Report results by slice: topic, risk level, language, customer segment. A release that improves general questions while getting worse on pricing or security answers is a regression, whatever the average says.

Make it a gate

Run the set automatically on every change: a new model version, a prompt edit, a new tool, new documents in the index. Set a pass threshold per slice before you see the results, and block the release when a slice falls below it. NIST’s AI Risk Management Framework puts this work under its Measure function: testing, evaluation, verification and validation, carried on after deployment rather than done once.4

Keep the set alive

  • Add every real failure from production as a new case, with its correct answer.
  • Retire cases that no longer reflect how people use the system.
  • Re-check model-based graders against people on a regular schedule.
  • Version the set with the code, so every result can be reproduced.

Open-source frameworks such as OpenAI’s Evals show the pattern in code.5 We design eval suites in AI Strategy & Agents and run them as release gates in AI Product & Automation.

Sources. Where the facts come from.

Numbered as they are cited in the text. Each link opens the original.

  1. Working with evals

    OpenAI API documentation

    Back to the text

  2. LLM01:2025 Prompt Injection

    OWASP Gen AI Security Project · 2025

    Back to the text

  3. AI RMF Core

    NIST Trustworthy and Responsible AI Resource Center

    Back to the text

  4. openai/evals

    OpenAI on GitHub

    Back to the text

Let’s build what happens next.

Tell us what you’re building. We’ll answer straight.

Book a discovery call

Three ways to start

  1. 01About 2 minutes

    A quick question

    You get A reply from a lead, not a sales queue

  2. 02About 8 minutesMost useful

    A project brief

    You get Options and a first scope after one call

  3. 03About 15 minutes

    A formal RFQ or RFP

    You get Receipt confirmed and a named bid lead

Every engagement starts with a written scope and a quote agreed before work begins. How each package is priced