Platform / Custom models / Evaluation and safety
Evaluation and safety for custom models: nothing is promoted until it has passed
How a custom model is tested on real domain tasks and on safety before it serves anyone, and why every adaptation is tested again from the start.
Safety first is a working rule, not a slogan: a custom model does not move towards production until it has passed domain evaluation, safety evaluation and a regression gate against the model it would replace. This page sets out what goes into each of those, how red-teaming fits, why fine-tuning can weaken safety behaviour, and what evidence a sign-off should rest on. It also says plainly what Swfte has built for this today and what is still on the roadmap.
Evaluation comes before promotion, every time
A model version moves through stages: development, staging, production. Evaluation is what decides whether it may move. That sounds obvious, but the common failure is the opposite order: a team ships a model because it looked good in a handful of demonstrations, then builds the evaluation after the first complaint. By then there is no baseline to compare against and no record of what was known at the time of release.
The rule here is simple. Before a candidate is promoted, it is run against the same suites as the model in production, on the same cases, and the results are kept. A candidate that does better on the domain task and no worse on safety and general behaviour can move on. One that does not stays where it is, whatever the deadline.
Domain evaluation built from real tasks
A domain suite is only as good as the cases in it. Build it from the work the model will actually do.
Real tasks
Cases drawn from the work itself: the questions people ask, the documents they classify, the drafts they write. Not invented puzzles that resemble the work.
Held-out cases
Cases the model has never been trained on, kept apart from training data and checked for near-duplicates. Without them, a score measures memory.
Expected outputs and graders
For each case, what a correct answer looks like and how it is judged: an exact match, a rubric applied by a person, or a checked rubric applied by a model.
Hard and rare cases
The edge cases reviewers worry about, over-represented on purpose, because an average score hides them.
Safety evaluation
Each of these is measured on the candidate and on the model it would replace.
Refusal and over-refusal
Does it decline what it should, and does it still answer legitimate questions in your domain? A model that refuses too much pushes people to tools you do not control.
Jailbreak and prompt injection
Attempts to talk the model out of its policy, and instructions hidden in retrieved content. Retrieved content is untrusted and must not override policy.
Harmful content
Categories of output your organisation will not produce, tested directly and through indirect requests.
Bias
Whether outputs change in ways they should not when names, roles or other attributes in a case are changed.
EU-language parity
Safety and quality measured in every language your users write in. Refusal behaviour is often weaker outside English.
Regression gates and red-teaming
A regression gate compares the candidate with the model currently in production on every suite. If the candidate is worse beyond the margin you set in advance, on any suite you have marked as blocking, promotion stops. The gate blocks; it is not a warning that someone can click past. Changing the margin is itself a reviewed decision, recorded with the reason.
Red-teaming adds people trying to break the model on purpose, with reference attack tooling alongside them. It finds failures that fixed suites miss, and every finding becomes a new case in the suite. Be exact about what a clean run means: no known attack succeeded during that run. It does not mean the model is safe in any absolute sense, and a report that says so should not be trusted.
Why safety can erode after fine-tuning
A base model's refusal behaviour is itself the product of training, and further training can disturb it. Published research has found that fine-tuning on entirely benign data can weaken refusals, and that a small number of harmful examples can weaken them a great deal. Adapters are not exempt. Quantising the weights for serving can shift behaviour too, most visibly outside English.
So every adaptation is treated as a new model. A new adapter, a new training run on refreshed data, a new quantisation of the same weights: each goes back through the full safety suite and the regression gate, and the artifact that is tested is the artifact that is served. Swfte calls this safety-first adaptation, and it is the design intent behind Swfte Safety, published at /models/swfte-safety as a design-intent model card. That card publishes no measured results, and we say so plainly.
The same method Swfte uses for open-weight models
The approach on this page is the one Swfte publishes for testing open-weight models at /open-source-model-testing: your own tasks, a safety suite and an EU-language suite, run against the candidate, with the result deciding whether it goes further. A custom model is held to the same method as any open-weight model you would deploy, plus the comparison with the model it replaces.
The named domain suites Swfte would provide are <named evaluation suites - founder to fill>. Until those exist, the suites are yours, built from your work, and the method is what keeps them honest. That is no hardship: a suite built from your own cases tells you more about your model than any public benchmark, and it stays with you whichever model you run next.
Sign-off and the evidence pack
Promotion to production is a decision a named person makes, on evidence that is kept.
- 01
Assemble the results
Domain, safety, multilingual and regression results for the candidate and the production model, on the same case versions.
- 02
Attach the lineage
The base model and revision, the dataset datasheet, the training configuration and the hash of the weights that were tested.
- 03
List the open findings
Red-team findings not yet fixed, known weaknesses and the mitigations at the gateway, written in plain language.
- 04
Record the decision
Who approved, at which stage, and on which evidence. The Model Vault keeps an audit log of promotions that can be exported.
What exists today, what is on the roadmap, and a worked example
Today you can score single-turn chat in Studio's evaluation, put guardrails and content policy in front of a model at the Connect gateway, with secret and personal-data detectors and redaction, and follow the published testing method. An agent evaluation sandbox with recorded, replayable traces is in progress as internal tooling. Named domain suites, an automated red-team pipeline and gated promotion wired to evaluation results are on the roadmap.
A worked example. A claims team has a candidate adapter that drafts summaries. It passes the domain suite clearly. On the safety suite it holds refusals in English but declines fewer harmful requests in two other EU languages than the production model does. The gate blocks promotion. The team adds multilingual safety examples to training, retrains, and runs everything again. The second candidate passes, a named reviewer signs the evidence pack, and the model moves to staging.
Evaluation and safety: built, in progress and roadmap
| Capability | Status | Notes |
|---|---|---|
| Single-turn evaluation in Studio | Built | Scores single-turn chat against your cases. |
| Content policy and redaction at the Connect gateway | Built | Secret and personal-data detectors with a redact action. |
| Published open-weight testing method | Built | At /open-source-model-testing. |
| Swfte Safety model card | Built | A design-intent card. No measured results are published. |
| Version stages and promotion audit log | Built | Development, staging, production and archived in the Model Vault. |
| Agent evaluation sandbox with replayable traces | In progress | Internal tooling, not a customer feature. |
| Named domain evaluation suites | Roadmap | Designed for. Not built. |
| Automated red-team pipeline | Roadmap | Manual red-teaming with reference tooling is the method today. |
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Where this fits in the loop
Evaluation and safety sit inside the custom model station: the model is hardened and proven here before any agent uses it.
- 01Company brainHolds what the organisation knows, with evidence statuses, history and access rules.
- 02Custom modelAdapted on data chosen from the brain, then evaluated and hardened before it ships.(this page)
- 03Governed agentsUse the model and read the brain, inside a Trust Profile, with approval where it matters.
- 04OutcomesWhat happened: approvals, corrections, results and cost, all on the record.
The four arrows
- Company brain to Custom model: select, sanitise, adaptRoadmap
Choose training data from the brain, remove what must not reach a model, adapt an open-weight base. The sanitisation gateway is in progress, and the data selection and training steps are on the roadmap.
- Custom model to Governed agents: serve, governBuilt
Serve the model on dedicated infrastructure behind the Connect gateway and bring agents onto it under policy. Model hosting and the gateway are built.
- Governed agents to Outcomes: act, recordBuilt
Agents act within their Trust Profile, with human approval for consequential steps, and every action is recorded.
- Outcomes to Company brain: written back as evidenceRoadmap
Outcomes return to the brain as new evidence with a status, and they decide when the model needs retraining. The write-back is on the roadmap.
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Frequently asked questions
What has to pass before a custom model is promoted?
Domain evaluation on held-out real tasks, safety evaluation covering refusal, over-refusal, jailbreaks, prompt injection, harmful content, bias and EU-language parity, and a regression gate against the model in production. A failing blocking suite stops promotion.
Does Swfte Safety have published benchmark results?
No. Swfte Safety is published as a design-intent model card. It sets out what the model is designed to do and how it is meant to be tested, and it publishes no measured results. We say that plainly rather than imply a score.
Why re-test a model after fine-tuning on harmless data?
Because refusal behaviour comes from training and further training can weaken it, even when the new data is benign. Quantisation can shift behaviour too. Every adapter, retrain or quantised artifact is treated as a new candidate and goes through the full safety suite again.
Does a clean red-team run mean the model is safe?
No. It means no known attack succeeded during that run. Red-teaming finds failures that fixed suites miss, and each finding becomes a new test case, but no run proves the absence of every failure.
What evaluation can we run in Swfte today?
Single-turn evaluation in Studio, guardrails and content policy at the Connect gateway, and the published method for testing open-weight models. An agent evaluation sandbox is in progress as internal tooling. Named domain suites and automated red-teaming are on the roadmap.
Does passing these tests make our model EU AI Act compliant?
Swfte does not claim that any model is compliant. Evaluation results, sign-offs and audit logs are technical evidence that helps an organisation meet its own obligations. The posture depends on use case, jurisdiction, deployment and configuration, and on your own legal assessment.
Who signs off a promotion?
A named person your organisation chooses, at each stage, on an evidence pack: results against the production model, the lineage of base model and data, open findings and mitigations. The Model Vault records promotions in an audit log that can be exported.
Take evaluation and safety further with Swfte
Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.