← The journal
Evaluation Guide

Evaluate an Open-Source LLM Before Production: Checklist

A checklist to evaluate an open-weight LLM before production: licence, supply chain, evals, safety and latency.

Swfte Journal / Evaluation Guide

Picking an open-weight model by leaderboard rank is how teams end up with a model that is excellent at a benchmark and wrong for the job. Benchmarks are useful for shortlisting and poor for deciding. The decision is a small, boring, repeatable evaluation on your own tasks, with your own safety requirements, in your own languages, on the hardware you will actually run, and with a record you can show to someone who asks.

This is the checklist we use as the basis of our open-source model testing methodology. It is ordered so that cheap, decisive checks come first: there is no point running a two-day evaluation on a model whose licence you cannot use.

Before you run anything: write down the decision

An evaluation without a decision rule is a demo. Before the first run, write down:

  • The job. What task, for whom, with what inputs and what a good output looks like.
  • The constraints. Data classification, region, hardware budget, latency target, languages, and who is allowed to approve the result.
  • The baseline. What you use today (a hosted API, a previous model, a rule-based system). Every score is a comparison with it.
  • The pass rule. For each suite, what result lets the candidate go forward and what stops it. Setting this after you have seen the numbers is how evaluations get bent to fit the model someone already likes.

Step 1: Licence and use rights

Do this first because it is free and decisive. Read the licence file and any acceptable use policy that ships with the exact checkpoint, not the licence of the family in general. Patterns differ widely:

  • Permissive. Apache 2.0 and MIT allow commercial use, modification and redistribution with attribution and a notice. Qwen's open-weight releases, Gemma 4 and most recent Mistral releases are Apache 2.0, and recent DeepSeek releases such as R1 are MIT.
  • Permissive with a scale clause. Kimi K2's modified MIT adds a requirement to display the model name in the product interface above 100 million monthly active users or 20 million US dollars in monthly revenue. Irrelevant for most, binding for some.
  • Custom community licences. Llama's licence is custom rather than OSI-approved, with a 700 million monthly-active-user clause, and the Llama 4 acceptable use policy withholds rights to its multimodal models from EU-domiciled individuals and EU-headquartered companies. An EU team needs counsel to read it before building on it.
  • Split licences and exceptions. Early DeepSeek V3 had MIT code and a separate model licence for the weights. Some specialist models from otherwise permissive makers ship under non-commercial terms. Some flagship "open" families have API-only flagships.

Record the licence text, the date you read it, and the commit hash of the checkpoint it applied to. Licences change between releases, which is why this belongs in the evaluation record rather than in someone's memory. This is orientation, not legal advice.

Step 2: Supply-chain check

You are about to run a few hundred gigabytes of files someone else produced. Treat that as you would any third-party dependency.

  1. Pin the revision. Download at a full 40-character commit hash, not a branch. Branches move.
  2. Prefer safetensors. Pickle-based checkpoints (.pt, .pth, .bin) can execute code when loaded. Safetensors stores raw tensor data and a JSON header, and an independent audit by Trail of Bits found no critical flaw leading to arbitrary code execution in it.
  3. Reject remote code. Loading flags such as trust_remote_code=True execute code from the repository. Leave it off unless you have reviewed that code, and prefer models that load with standard architectures.
  4. Verify hashes and keep a manifest. Record a hash for every file and compare on every move between environments. Where the publisher supports it, verify a signature; the OpenSSF Model Signing specification defines a detached signature over weights, configs and tokenizers as one unit.
  5. Scan anything that is pickle. Scanners such as picklescan help, but they use deny-lists and have been bypassed in the wild (ReversingLabs documented malicious models that evaded Hugging Face's scanner in early 2025 by using a different compression wrapper). Scanning is a filter, not a guarantee.

We go deeper in model supply-chain security.

Step 3: Capability on your tasks

Public benchmarks tell you the model is in the right league. Your tasks tell you whether it works.

  • Build a held-out set of 100 to 300 real cases per task if you can, drawn from actual inputs with sensitive content handled appropriately. Keep it out of any prompt tuning or fine-tuning.
  • Score with the right tool. Exact match or schema validity for structured output; reference-based checks where there is a right answer; a calibrated judge model or human review where there is not, with the judge validated against human labels on a sample.
  • Test the things that fail in production: long-context retrieval, tool-calling reliability (does it produce valid arguments, does it call the right tool), structured-output validity under load, and instruction-following when the system prompt is long.
  • Fix the run conditions. Same prompt template, same sampling settings, same chat template. Open-weight models are sensitive to the chat template, and a wrong one quietly costs quality.
  • Report sample size. A difference of two points on 50 cases is noise.

For public comparability, a harness such as EleutherAI's lm-evaluation-harness or the UK AI Security Institute's Inspect framework gives reproducible runs. Our own public data and method live on the benchmarks pages and the leaderboard.

Step 4: Safety

Safety evaluation should be at least as large as capability evaluation for anything that will touch customers or act through tools.

  • Harmful-request refusal. Does it decline clearly harmful requests? Use a curated set; HarmBench provides a standardized one and reports attack success rate.
  • Over-refusal. Does it wrongly decline benign requests that merely resemble harmful ones? XSTest-style pairs measure this. A model that refuses everything passes the first test and is useless.
  • Jailbreak resistance. Role-play, encoding tricks, multi-turn escalation. Automated scanners such as NVIDIA's garak, Microsoft's PyRIT and promptfoo run large attack corpora, and manual red-teaming finds what the corpora do not. See the red-teaming checklist.
  • Prompt injection. Instructions hidden in documents, tool output and web content. This is the top-ranked risk in the OWASP Top 10 for LLM Applications 2025, which also lists supply chain (LLM03) and sensitive information disclosure (LLM02).
  • Bias and fairness. Whether outcomes shift with a name, gender or nationality in otherwise identical prompts, for any use that affects people.
  • Leakage. Can it be coaxed into revealing the system prompt or sensitive context?

Remember what a passing result means: the attacks you ran did not succeed. It does not mean there are none.

Step 5: Languages

If you serve users outside English, run the safety and capability suites in each language and compare with English. Safety behavior often degrades in lower-resource languages, which makes language coverage a risk topic and not only a quality one. Use native or professionally translated cases rather than machine-translated English, because the latter hides exactly the failures you are looking for. For an EU deployment, set a parity limit per language and write it down in advance.

Step 6: Latency, throughput and cost

Measure on the target hardware and serving stack, under realistic concurrency, with realistic prompt and output lengths.

  • Time to first token and tokens per second per request, at the concurrency you expect.
  • Throughput per GPU at that quality target.
  • Memory fit. Weights are roughly two bytes per parameter in bf16, one in FP8 and half a byte at 4-bit, plus KV cache and headroom. Mixture-of-experts models need all experts resident even though few are active per token.
  • Cost per completed task, not per million tokens. A verbose model can be cheaper per token and more expensive per task.

If you plan to quantize, evaluate the quantized artifact, not the original. Quantization changes the model, and refusal behavior and non-English quality can shift. The GPU reference lists per-card memory and bandwidth for the sums, and the self-hosted inference guide covers the engineering.

Step 7: Regression gates

Everything above is a one-off unless you make it a gate. A gate compares a candidate with the revision currently in production on pinned suites and blocks promotion when behavior regresses beyond a tolerance you set in advance. Apply it to more than new models:

  • a new quantization,
  • a serving-engine upgrade,
  • a new system prompt,
  • a new fine-tune or adapter,
  • a new base revision.

Each of these changes behavior. Keep the previous approved revision deployed and routable until the new one has run on live traffic, so rollback is a configuration change. Put the endpoint behind a gateway such as Connect so routing, approved-model policy and logging are applied to every request.

Step 8: Record and decide

The output of the evaluation is a short record, not a slide. For each candidate keep: model and exact commit hash; licence text and date read; hash manifest; each suite's name, harness version, prompt template, sampling settings, sample size and run date; results against the baseline and against the pass rule; reviewer and decision. That is what lets you answer "what is running and why was it approved" months later, and it is the kind of evidence an audit asks for. Only publish a result you can reproduce, with its method.

A one-page checklist

  1. Decision, baseline and pass rules written before any run.
  2. Licence and use rights read for the exact checkpoint, including regional clauses.
  3. Revision pinned; safetensors only; no remote code; hashes recorded.
  4. Capability suite on your own held-out tasks, with fixed run conditions.
  5. Safety suite: refusal, over-refusal, jailbreak, injection, bias, leakage.
  6. Language suites with parity limits.
  7. Latency, memory fit and cost per completed task on target hardware.
  8. Regression gate against the production revision, for every change type.
  9. Record kept; rollback path tested.

If you want this as a managed loop on dedicated, in-region infrastructure, that is what the deploy models hub describes: Pick, Evaluate, Harden, Deploy, Monitor. Related reading: deploying an open-source LLM, best open-weight LLMs of 2026, and how to fine-tune an LLM on your own data.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.