Validating Your AI in Ten Steps: Evals to Evidence Pack
A ten-step map for validating your AI: eval sets, judge calibration, safety tests, regression gates and evidence.
Our guide, how to validate your AI, runs to ten steps, with a first runnable evaluation, a judge-calibration step and a CI gate. This post is the map: what each step is for, the mistake it prevents, and how the pieces fit together. Use it to decide where you are and which part of the full guide to open first.
The guide's one-line answer: validating an AI system means proving, on cases that look like your real work, that it meets thresholds you set before testing, and then proving it again every time something changes.
Why validation is a loop, not a launch gate
A model, a prompt, a retrieval index, a tool definition and a model provider's silent update can each change behaviour. A one-off test before launch describes the system on one day. Validation that works treats the thresholds, the cases and the results as artefacts that live alongside the code and run on every change. That is also what makes the evidence credible later, when a customer, an auditor or a regulator asks how you know it works.
The ten steps
1. Define the task, the decision and the risk tier. Write down what the system does, what decision it informs or makes, who is affected by a wrong answer, and how bad that is. The risk tier decides how strict everything after this is. A drafting aid and a system that influences access to credit are not validated to the same standard. For regulated settings, our EU AI Act high-risk checklist helps you see which duties may apply.
2. Build the eval set from real cases. Real inputs, with the outputs you would accept. Sample across the range of work, not only the easy middle, and include the awkward cases: ambiguous inputs, other languages, empty fields, long inputs. Keep a held-out portion the system and its builders never tune against, or the results flatter you. The guide says plainly that the right size is a judgement call that depends on how precisely you need to measure, not a number to copy.
3. Choose metrics for the task type and set thresholds. Classification, extraction, summarisation, retrieval and agent tasks need different measures. Where an answer can be checked by code (a field matches, a JSON schema is valid, a tool was called with the right arguments), use deterministic checks first. They are cheap, repeatable and do not argue. Set the pass thresholds now, before you see results. Thresholds chosen after the fact get bent to fit the system you already like.
4. Run a first eval locally with promptfoo. The guide uses an open evaluation tool with a configuration file, so you get a working loop in an afternoon and can see what a result looks like. The point of this step is to have something running, not something perfect.
5. Calibrate any LLM judge against human labels. A language model can grade open-ended outputs at scale, but it has known biases, such as favouring longer answers, favouring answers in a particular position, and favouring its own style. The guide tells you to have people label a sample, compare the judge's scores with theirs, and only trust the judge where they agree. Report that agreement alongside the scores. A judge you have not calibrated is an opinion with decimal places.
6. Add safety and abuse tests. Include cases where the system should refuse, should not leak, or should not follow instructions hidden in content. Our guides on red teaming an LLM and stopping prompt injection go deeper, and the AI red teaming page describes the approach.
7. Fail the build on regressions. Run the evals in continuous integration. Compare each run with the last accepted baseline, and fail when a metric drops past the tolerance you set. Without this step the eval set becomes a report nobody reads. With it, a prompt change that quietly makes one category worse is caught the day it is made.
8. Sample live traffic for human review. Pre-release cases never cover everything real users do. Draw a regular sample of production interactions, have a person grade them, and add the failures to the eval set. This is how the set stays honest as the world changes.
9. Monitor production against the pre-release numbers. Track the same measures you used before launch, plus cost and latency, and alert on change. Our guide to monitoring AI agents in production covers the instrumentation.
10. Assemble the evidence pack. Keep, in one place and under version control: the task and risk statement, the eval set and its provenance, the thresholds and when they were set, the results per version, the judge calibration, the safety test results, the review samples, and the decisions taken. This is what you hand over when asked. The platform governance page and the evaluation and safety page describe how Swfte frames this work.
A worked picture
Suppose you run an assistant that drafts replies to customer emails for a person to send. The decision it informs is what the customer is told, the risk is moderate (a wrong reply costs goodwill and sometimes money), and the thresholds might say that no reply may state a refund that policy does not allow, and that drafts must be judged usable by reviewers at least as often as the current process.
The eval set is a few hundred real emails, sampled across categories and languages, each with a reply a senior agent accepted. Deterministic checks catch the refund rule and the required fields. A calibrated judge scores tone and completeness, and you report how closely it agrees with two people who graded the same sample. The safety cases include an email that contains instructions aimed at the assistant. CI runs the set on every prompt change and fails if the refund-rule check drops at all or the usability score drops past tolerance. In production, someone grades a sample of drafts each week, and the failures join the set. When someone asks how you know it works, you open the evidence pack.
None of that needs a platform. It needs discipline about the order: risk, cases, thresholds, tests, gate, review, record.
Where people go wrong
- Testing on the examples they tuned on. The numbers look excellent and mean nothing.
- Averaging away the failures. A system that is right 95% of the time and catastrophically wrong on one customer segment has a segment problem, not a 95% problem. Report results by slice.
- Treating a public benchmark as validation. Benchmarks help you shortlist. They do not tell you how the system does your work. Our open-source model testing methodology uses them only for that.
- No baseline. Without the previous version's scores, you cannot tell improvement from drift.
- Letting the judge grade its own family. If the same model writes and grades, expect inflated scores.
- Writing thresholds after the results.
What it costs, and how to keep it small
Teams worry that validation is a large project. It does not have to be. A first useful version is a spreadsheet of fifty real cases, a script that runs them and a person who reads the failures. That already beats most launches. The cost grows with the risk: a low-risk drafting aid needs a small set and a light review sample, while a system affecting people's access to services needs a larger set, sliced results, stricter thresholds and documented human review. Scale the effort to the tier you set in step 1, and write down why you chose that scale.
Two habits keep the cost down. First, add every production failure to the set the day you find it, so the set grows from real incidents rather than imagination. Second, keep the harness boring: plain files in version control, one command to run it, and results stored per version. Nobody maintains a clever framework for long.
Which step to open first
If you have nothing: start at step 1 and 2, and have the eval set before you have the tooling. If you have evals but no gate: step 7. If you have a gate but no live review: step 8. If you are being asked for proof: step 10, and work backwards to see what is missing.
Related
- How to evaluate an open-source LLM for choosing a model.
- How to choose an LLM for your company for the decision process around it.
- How to audit AI systems for the external view of the same evidence.
The full guide has the commands, the configuration, expected output and a troubleshooting table, with a verification date and sources on every step. You can do all of it with open tools and without Swfte.