LLM Evaluation: How to Test a Model Before and After Launch
A practical guide to LLM evaluation: test sets, metrics by task, judge models, regression gates and what to log.
LLM evaluation is the practice of testing a model or an AI feature against a fixed set of cases drawn from your own task, scoring each answer with a rule written down in advance, and repeating the run every time the model, the prompt or the data changes. A one-off benchmark score does not do this job. A test set you own, a grader you trust and a gate in your release process do. Last verified 2026-10-07 against the vendor documentation and papers listed under Sources.
This guide covers what to test before launch, what to watch after it, how to choose a metric for each type of task, how far an LLM judge can be trusted, and what to log so a failure can be reproduced.
What is the difference between offline and online evaluation?
Offline evaluation runs a fixed test set against a candidate before anything reaches a user. It is repeatable: the same inputs, the same graders, a comparable score. Use it to choose between models, to compare two prompts and to stop a regression from shipping.
Online evaluation looks at real traffic after launch. It catches what a fixed set cannot: new kinds of input, drift in how people use the feature, and failures that only appear at volume. Typical signals are user ratings, edits and rejections, escalations to a person, and a sampled review of production answers.
You need both, and they feed each other. Every failure found online should become a new offline case. OpenAI's evaluation guide frames this as eval-driven development: "Evaluate early and often. Write scoped tests at every stage." The same page warns against "vibe-based evals", where "it seems like it's working" is the only evidence.
How do you build a test set from real tasks?
Start with the work, not with a public benchmark. Anthropic's documentation says to "design evals that mirror your real-world task distribution" and not to forget edge cases such as irrelevant input, overly long input, poor user input and ambiguous cases.
A workable method:
- Collect real inputs from logs, tickets or the people who do the task today. Remove personal data before storing them.
- Write the expected output or the grading rule for each one. A single correct answer where one exists, a rubric where it does not.
- Over-represent the hard and rare cases, because an average score hides them.
- Keep a held-out slice that is never used to tune prompts, so the score measures the system and not your memory of the test.
- Version the set. A changed test set changes the baseline, so record the version next to every score.
Anthropic also states a preference that matters for budgets: "More questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals." Volume beats polish, provided the grader is checked against people (see below).
Which metric fits which task?
Pick the simplest grader that can still tell a good answer from a bad one. The table follows the grading methods in Anthropic's and OpenAI's documentation and the metrics reported in the HELM paper.
| Task type | Grader | Use when | Watch for |
|---|---|---|---|
| Classification, extraction, short factual answers | Exact match or string check, after normalising case and whitespace | There is one correct answer | Correct answers phrased differently score as failures |
| Summaries and rewrites | Overlap metrics such as ROUGE-L, or embedding similarity | You hold reference texts | High overlap does not prove the summary is faithful |
| Open-ended answers, tone, helpfulness | Rubric scoring by a person or a judge model | No single right answer exists | Rubric wording drives the score; write each score level out |
| Retrieval-augmented generation (RAG) | Retrieval metrics (was the right passage retrieved?) plus answer metrics (is the answer supported by it?) | An answer depends on your documents | A good answer can hide a bad retrieval step; score the two stages separately |
| Code and structured output | Run it: tests pass, schema validates | Output can be executed or parsed | Passing tests can still miss the intent |
| Safety and policy | Binary classifier or rubric over adversarial and ordinary prompts | The feature handles sensitive content | Over-refusal: a model that refuses everything scores well on harm |
The HELM paper from Stanford (Liang and colleagues, published in Transactions on Machine Learning Research in 2023) is a useful reminder that one number is never enough. It measured accuracy, calibration, fairness, bias, toxicity, efficiency and resilience to input changes across 42 scenarios for 30 models. You do not need that breadth. You do need to decide which two or three dimensions matter for your feature and track each separately.
Can a model grade another model's answers?
Yes, with checks. LLM-as-a-judge means asking a strong model to score or rank answers against a rubric. The foundational paper, "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (Zheng and colleagues, submitted June 2023), reports that strong judges such as GPT-4 reached over 80% agreement with human preferences, about the level at which humans agree with each other. The same paper names the limits: position bias, verbosity bias, self-enhancement bias and limited reasoning ability.
Those limits translate into working rules:
- Randomise answer order when comparing two answers, or score each answer alone, to reduce position bias.
- Do not reward length. Put length expectations in the rubric, and check whether longer answers score higher for no reason.
- Use a different model as judge from the one that wrote the answer. Anthropic's documentation says this is generally best practice, and it addresses self-enhancement bias.
- Write the rubric out. Both Anthropic and OpenAI show judge prompts with a fixed scale, an explicit definition for each score and an instruction to output only the score.
- Calibrate against people. OpenAI's guidance is to scale up once the judge reaches agreement with human annotations. Label a sample by hand, compare, and re-label whenever the rubric changes.
The 80% figure comes from chat-assistant benchmarks. It is not a promise about your task. Measure agreement on your own cases.
Where does human review still matter?
Use people for three jobs: writing the first rubric, calibrating the judge, and reviewing the cases where the judge and the rule disagree. Two reviewers per sampled case, with a way to resolve disagreement, gives a more reliable label than one. Do not use human review as the main scoring route at scale. It is slow and it does not repeat exactly.
How do you turn evaluation into a regression gate?
A gate is a rule in your release process: a candidate that scores worse than the current version beyond a margin you set in advance does not ship. Four details make it work.
- Fix the baseline. Score the production version on the same test-set version and store it.
- Set the margin before you see the result. Choosing a threshold afterwards turns the gate into a formality.
- Mark suites as blocking or advisory. Safety and critical-task suites block. Style suites can warn.
- Run it in CI on every change to the prompt, the model name, the retrieval settings or the tool definitions, and on a schedule, to catch changes you did not make.
OpenAI's guide for continuous evaluation says the same in fewer words: run automated tests on every change, watch for non-determinism, and grow the test set over time. Because outputs vary between runs, run each case more than once where the cost allows and compare distributions, not single values.
A note on tooling. OpenAI's evals guide carried this notice when read on 2026-10-07: the Evals platform is being deprecated, becomes read-only for existing users on 31 October 2026 and is scheduled to shut down on 30 November 2026. If you built on it, export your datasets and graders now. Open-source options include Hugging Face's Lighteval, which supports several inference backends, custom tasks and metrics, and sample-by-sample results.
How should you evaluate safety?
Treat safety as its own suite with its own gate. Anthropic lists privacy preservation among its example success criteria and states that even "hazy" topics such as ethics and safety can be quantified. In practice a safety suite combines:
- prompts the model must decline, and ordinary prompts it must still answer, so you measure over-refusal as well as harm;
- prompt-injection attempts hidden in retrieved documents and tool output;
- personal-data and secret leakage checks on outputs;
- the same cases in every language your users write in, because refusal behaviour often differs outside English.
A clean run means that no known attack succeeded during that run. It does not mean the system is safe. Add every attack that works to the suite.
What about cost and latency?
Anthropic's list of success criteria includes latency and price next to task fidelity, which is the right framing: a quality gain that doubles cost per request is a decision, not a free win. Record tokens in and out, time to first token, total latency and cost for every evaluation run, and set budgets per feature. Compare candidates on the three together. Report latency as a distribution, not an average.
What should you log?
Log enough to replay any answer later. For every request, in evaluation and in production:
- the exact input, the prompt template version, the model name and version, and the settings (temperature, tool list, retrieval parameters);
- the retrieved passages, with identifiers, and every tool call with its result;
- the output, the grader's score and the grader's reasoning, with the rubric version;
- token counts, latency and cost;
- the user's action afterwards: accepted, edited, rejected, escalated.
Apply the same data rules as to any other store of user content: redact secrets and personal data before logging, set a retention period and restrict who can read the logs. See LLM observability for tracing and cost per call, and AI observability for the platform view.
Where does Swfte fit?
Swfte gives you parts of this, not a complete evaluation product, and you do not need Swfte to run the method above.
- Studio has an evaluation endpoint,
/v2/evaluation, that scores single-turn chat. Status: Built. It does not score multi-turn conversations or agent runs. - The method for custom models, with domain suites, a safety suite and a regression gate against the model in production, is set out on evaluation and safety for custom models, which states what is built today. The method for open-weight models is published at open-source model testing.
- A sandbox for evaluating agent scenarios exists as internal tooling. Status: In progress. It is not a customer feature.
- Swfte does not publish named domain evaluation suites or benchmark scores for its own models, and this post does not either.
For step-by-step guidance read how to evaluate an open-source LLM and how to validate your AI. Swfte does not claim certification or compliance for the systems you test with any of this. See Trust for what Swfte states about its own platform.
Sources and last verified
Last verified 2026-10-07. Every dated or technical fact in this post was read from the pages below on that date. Anything that could not be confirmed is left out or marked as not verified.
- Anthropic: create strong empirical evaluations. Success criteria, task-specific evals, grading methods (exact match, similarity, ROUGE-L, LLM-based scales) and the advice to use a different model as judge.
- OpenAI: working with evals. Definition of evals and the notice that the Evals platform becomes read-only on 31 October 2026 and shuts down on 30 November 2026.
- OpenAI: evaluation best practices. Eval-driven development, vibe-based evals as an anti-pattern, judge biases, calibration against human labels and continuous evaluation.
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685). Over 80% agreement with human preferences for strong judges, and the named position, verbosity and self-enhancement biases.
- Liang et al., Holistic Evaluation of Language Models (arXiv 2211.09110). Metric dimensions (accuracy, calibration, fairness, bias, toxicity, efficiency and more) across 42 scenarios and 30 models, used to show why one number is not enough.
- Hugging Face: Lighteval documentation. Open-source evaluation toolkit: multiple backends, custom tasks and metrics, sample-by-sample results.
Frequently asked questions
What is LLM evaluation?
LLM evaluation is testing a model or an AI feature against a fixed set of cases, scoring each answer with a rule written in advance, and repeating the run whenever the model, prompt or data changes. It differs from a public benchmark because the cases come from your own task. The result is a score you can compare between versions, which lets you block a release that gets worse.
How many test cases do you need?
There is no verified minimum, and the vendor documentation read for this post does not give one. A practical start is a few dozen real cases that cover the common paths and the known hard ones, then growth with every failure found in production. Anthropic advises that more questions with slightly lower signal automated grading beat fewer questions with hand-graded quality, so favour volume once the grader is checked.
Can an LLM judge replace human reviewers?
Not entirely. The MT-Bench paper reports that strong judges such as GPT-4 reached over 80% agreement with human preferences on chat benchmarks, about the level at which humans agree with each other, but it also names position, verbosity and self-enhancement biases. Use people to write the rubric, to calibrate the judge on a labelled sample and to review disagreements, and let the judge do the volume.
How often should you run evaluations?
Run them on every change to a prompt, model name, retrieval setting or tool definition, and again on a schedule. OpenAI describes evaluation as continuous: automated tests on every change, attention to non-determinism, and a test set that grows over time. Because outputs vary between runs, repeat each case where the cost allows and compare distributions rather than single scores.
What is the difference between a benchmark and your own evaluation?
A benchmark measures general ability on shared tasks, and your own evaluation measures whether the system does your job. Public benchmarks help to shortlist models, but Anthropic advises designing evals that mirror your real-world task distribution, including edge cases. The HELM paper shows how many dimensions a benchmark can cover, such as accuracy, calibration and toxicity. Choose the two or three that matter for your feature.
Does Swfte offer LLM evaluation?
Partly. Studio has an evaluation endpoint that scores single-turn chat, which is Built, and Swfte publishes a testing method for open-weight and custom models. A sandbox for agent scenarios is internal tooling that is In progress, not a customer feature. Swfte does not publish named domain evaluation suites or benchmark scores for its own models. You can run the method in this post with any tooling.