Evaluation-driven
How Swfte tests open-source models before they reach customers
Capability, safety, multilingual, latency, licence and supply-chain tests, with regression gates and a policy on what we publish.
An open-weight model arrives with a leaderboard score and a licence file. Neither tells you whether it will refuse a harmful request in Polish, resist an instruction hidden in a PDF, or ship with weights that were tampered with on the way. Our position is evaluation-driven: no model is offered to customers because it is popular, and none is promoted because it scored well once. This page is the methodology. It also says plainly what we have not published yet.
The pipeline, in order
Every candidate runs the same gates in the same order. A failure stops the run.
Intake
Record the exact repository, full commit hash, licence file, model card and the claimed capabilities. Nothing is tested that cannot be re-fetched.
Supply-chain check
Verify file hashes, require safetensors, scan any pickle artifact, and reject models that need remote code execution to load.
Licence review
Read the licence and acceptable use policy for commercial use, redistribution, scale clauses, attribution and regional restrictions.
Capability evals
Run the model on tasks that look like customer work, plus public benchmarks for comparability. Record harness, prompt template, sampling settings and date.
Safety evals
Test refusal of harmful requests, over-refusal of benign ones, jailbreak and prompt-injection resilience, harmful content and bias.
Multilingual evals
Repeat the safety and capability suites in EU languages and compare with English.
Latency and cost
Measure time to first token, throughput and cost per completed task on the target hardware and serving stack.
Regression gate
Compare against the revision in production on the pinned suites. A regression blocks promotion.
Publish
Publish what we can reproduce, with the method, and say what we did not run.
What each test area covers
Capability
Task accuracy on held-out cases drawn from the work customers actually do, long-context behavior, tool-calling reliability and structured-output validity. Public benchmarks are included for comparability, but benchmarks leak into training data over time, so private cases carry the weight.
Safety: refusal and over-refusal
Refusal of clearly harmful requests, and the opposite failure, refusing benign requests that merely sound risky. Both are measured because a model that refuses everything passes the first test and fails the second.
Safety: jailbreak and prompt injection
Direct jailbreaks, and indirect injection through documents, tool output and web content, which is the top-ranked risk in the OWASP Top 10 for LLM Applications. We also look at system-prompt leakage and sensitive-information disclosure.
Safety: harmful content and bias
Harmful-content categories, plus bias and fairness probes relevant to the use case, such as whether outcomes shift with a name, gender or nationality in an otherwise identical prompt.
Multilingual (EU languages)
The same suites in the EU languages customers use, compared with English, so a gap in a lower-resource language shows up before a user finds it.
Latency and cost
Time to first token, tokens per second under concurrency and cost per completed task, on the target hardware. A verbose model can be cheaper per token and dearer per task.
Licence review
Commercial use, scale clauses, attribution, redistribution of derivatives and regional restrictions. Licences differ per checkpoint and change between releases.
Supply chain
Hash verification against a pinned revision, safetensors only, scanning of any pickle file (knowing that scanners have been bypassed), remote code execution off, and a record of provenance.
Regression gates
| Gate | What it checks | Blocks promotion when |
|---|---|---|
| Supply chain | Hashes match the pinned revision; safetensors; no remote code. | A hash differs, an unsafe artifact is present, or remote code is required. |
| Licence | Licence and acceptable use policy permit the intended deployment. | Use or region is excluded, or a clause cannot be met. |
| Safety | Refusal, over-refusal, injection, harmful-content and bias suites against the production revision. | Safety behavior regresses beyond the tolerance set for that suite: <regression tolerances per suite — founder to fill>. |
| Capability | Customer-task suite against the production revision. | Task quality regresses beyond tolerance. |
| Multilingual | EU-language suites, compared with English. | A supported language falls behind the agreed parity limit: <parity limits per language — founder to fill>. |
| Cost and latency | Time to first token and cost per completed task on target hardware. | The model cannot meet the route’s latency or cost budget. |
Quantizations, serving-image upgrades and system-prompt changes go through the same gates as new models, because each of them changes behavior.
Reference tooling and standards
The method is built on open, reproducible tooling of the kind the field uses: EleutherAI’s lm-evaluation-harness and the UK AI Security Institute’s Inspect for capability and behavior evaluation; NVIDIA’s garak, Microsoft’s PyRIT and promptfoo for probing and red-teaming; HarmBench for attack-success measurement and XSTest for exaggerated refusal; and the OWASP Top 10 for LLM Applications as a risk taxonomy. These are named as reference implementations. Which harness runs in a given suite, and at which version, is recorded in that suite’s run record. <harness versions per suite — founder to fill>.
A clean run means no known attack succeeded, not that the model is safe. Red-team tools cover known attack corpora; they do not find the attack nobody has written yet, so we also keep manual red-teaming on the plan for models headed into sensitive workflows.
Publish-results policy
We publish a result only when we can reproduce it, and we publish it with the model revision, the harness and version, the prompt template, the sampling settings, the sample size and the run date. We report sample sizes next to scores, because small public samples produce wide confidence intervals.
We do not publish a number we cannot stand behind, and we do not fill gaps with figures from a model card or a press release. Where a result has not been published, the test log below says so. Our existing capability methodology and public data live on the benchmarks pages.
Test log
Open-weight models in our model directory, newest first. Directory facts are shown as listed; Swfte’s own test dates and verdicts are open fields until results are published.
| Model | Maker | Released | Licence (as listed in the model directory) | Context window | Directory row last verified | Swfte test date | Swfte verdict |
|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek | 2026-09-10 | MIT | 1M tokens | 2026-09-19 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| GLM-5.3 Flash | Z.ai (Zhipu AI) | 2026-08-26 | MIT | 1M tokens | 2026-09-19 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| GLM-5.3 | Z.ai (Zhipu AI) | 2026-08-18 | GLM-5.3 License | 1M tokens | 2026-09-19 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| Qwen3.8 27B | Alibaba Cloud | 2026-08-14 | Apache 2.0 | 262K tokens | 2026-08-17 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| DeepSeek V4 Pro | DeepSeek | 2026-08-13 | MIT | 1M tokens | 2026-08-17 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| Nemotron 3.5 Lightning | NVIDIA | 2026-08-11 | OpenMDW-1.1 | 1M tokens | 2026-08-17 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| Ling-3.0-Flash | Ant Group | 2026-08-05 | MIT | 256K tokens | 2026-08-17 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| Qwen3.8 Max | Alibaba Cloud | 2026-08-03 | Custom (revenue-share; not Apache 2.0) | 1M tokens | 2026-08-17 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| Kimi K3 | Moonshot AI | 2026-07-16 | Modified MIT | 1.0M tokens | Not recorded | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| Hunyuan Hy3 | Tencent | 2026-07-06 | Apache 2.0 | 256K tokens | 2026-08-17 | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| GLM-5.2 | Z.ai (Zhipu AI) | 2026-06-13 | MIT | 1M tokens | Not recorded | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
| MiniMax M3 | MiniMax | 2026-06-01 | Open weight (modified MIT expected) | 1M tokens | Not recorded | <Swfte test run date — founder to fill> | <Swfte verdict — founder to fill> |
The first six columns are read from Swfte’s model directory (src/data/ai-directory) and describe the model as listed there, not our test results. Dates and verdicts of Swfte’s own test runs are not published yet; each is a placeholder, not a pass.
Public leaderboard data lives at /benchmarks/leaderboard and /models/leaderboard.
Frequently asked questions
Which open-source models does Swfte test?
Candidates for deployment through Swfte, including the Llama, Qwen, DeepSeek, Gemma, Kimi and Mistral families, and customer-supplied or fine-tuned models. The test log lists the open-weight models in our model directory and the status of our own runs against each.
Do you publish test results?
Only results we can reproduce, with method, revision and date. Entries not yet published are marked in the test log rather than filled with borrowed numbers.
What does a regression gate do?
It compares a candidate with the revision currently in production on pinned suites and blocks promotion when behavior regresses beyond the tolerance for that suite.
Why test over-refusal as well as refusal?
A model that refuses everything passes a harmful-request test and is useless. XSTest-style probes measure whether benign requests that merely resemble harmful ones are wrongly refused.
Does a passing result mean a model is safe?
No. It means the known attacks and cases we ran did not succeed. Safety evaluation lowers risk and produces evidence; it does not prove absence of risk.
Do you test quantized models separately?
Yes. A quantized artifact is a different artifact and goes through the same gates.