Swfte model
Swfte Safety: a safety-first fine-tuned model for governed AI
A model fine-tuned to refuse, ask and escalate correctly, follow policy, and work across EU languages, so that governed agents have a safer default underneath them.
Most models are tuned to be helpful first and restrained second. For agents that act on real systems, the order matters the other way round: a model that does what it is told by whoever is talking to it is a liability inside a workflow. Swfte Safety is Swfte’s own fine-tuned model, designed safety-first. This page states what it is designed to do, what is still to be published, and what is not claimed. We separate design intent from measured results on purpose, and no measured results are published yet.
Design intent
What the model is built to do. These are intentions, not measurements.
Safety-first fine-tuning
Safety behavior is a training objective from the start, not a filter added after the fact: refusal, calibrated caution and policy adherence are trained and then evaluated, rather than assumed.
Refusal and escalation
Beyond yes and no, the model is meant to ask a clarifying question, decline, or request human approval. Those map to the platform’s policy verbs: allow, deny, warn, filter, escalate and require approval.
Policy-following
Meant to prefer a system-level policy over conflicting instructions in the user turn or in retrieved content, so an agent’s trust profile still means something when the model is under pressure.
Injection-aware reading
Meant to treat documents, tool results and web pages as untrusted data, which is the main defense against indirect prompt injection.
EU-language coverage
Meant to keep refusal and policy behavior consistent across the EU languages it supports, since safety behavior often degrades outside English.
Auditable output
Meant to produce structured, attributable output (what it decided, which policy it applied, what it needs approved) so the audit trail has something to record.
Design intent versus measured results
Left column: what the model is designed to do. Right column: what has been measured. Nothing has been measured and published yet, so every cell on the right is an open field.
| Property | Design intent | Measured |
|---|---|---|
| Refusal of clearly harmful requests | Decline, briefly and without lecturing, and offer a safe alternative where one exists. | <harmful-request refusal result — founder to fill> |
| Over-refusal of benign requests | Answer legitimate requests that merely resemble harmful ones, such as a security analyst asking how an exploit works. | <over-refusal result — founder to fill> |
| Prompt-injection resilience | Treat retrieved documents, tool output and web content as data, not as instructions, and flag attempts to override policy. | <prompt-injection resilience result — founder to fill> |
| Policy-following | Follow a system-level policy set (allowed actions, restricted actions, approval rules) over conflicting user instructions. | <policy-adherence result — founder to fill> |
| Escalation behavior | Request human approval or hand off when an action is out of policy, ambiguous or high-impact, instead of guessing. | <escalation-correctness result — founder to fill> |
| EU-language coverage | Keep refusal and policy behavior consistent across the EU languages it supports, not only English. | <per-language parity result — founder to fill> |
| Latency and cost | Small enough to run on dedicated infrastructure at practical cost. | <latency and cost result — founder to fill> |
Model card
Written in the style of our model cards. As of 2026-10-06, the base model, data provenance, results, release date, licence and availability are not published. They appear as placeholders so nobody mistakes them for commitments.
- Name
- Swfte Safety (working name; may change before release)Design intent
- Status
- Design intent published. No measured results published yet.Design intent
- Base model
- <base model and revision — founder to fill>To be published
- Parameter count and architecture
- <parameters and architecture — founder to fill>To be published
- Fine-tuning method
- <fine-tuning method — founder to fill>To be published
- Training data provenance
- <training data sources, licences and filtering — founder to fill>To be published
- Evaluation results
- <evaluation results and run dates — founder to fill>Measured
- Languages (measured)
- <measured per-language results — founder to fill>Measured
- Release date
- <release date — founder to fill>To be published
- Licence
- <licence — founder to fill>To be published
- Availability
- <availability: where and how it can be run — founder to fill>To be published
Intended use
- Governed agents that act under a policy set and a trust profile, where declining, asking or escalating is a normal outcome rather than a failure.
- SecOps triage: reading alerts and tickets, summarizing, proposing next steps, and handing consequential actions to a human.
- Regulated workflows where an auditable, policy-following assistant is worth more than a maximally permissive one.
Out of scope
- Sole decision-maker for legal, medical, financial or employment outcomes without human review.
- Any use that is prohibited under applicable law, including practices the EU AI Act prohibits.
- A substitute for runtime guardrails, access control, human approval or an evaluation gate. It is one layer in a defense, never the whole defense.
How it fits the platform
The model is one layer. In a Swfte deployment it runs on dedicated infrastructure behind the Connect gateway, under the approved-model policy; Nexus gives the agents that call it an identity, permissions and an audit trail; runtime guardrails check inputs and outputs; and human approval gates consequential actions. A safety-tuned model makes the lower layers easier to trust. It does not make them optional.
The same evaluation gate that applies to every open-source model applies to this one, including after each fine-tuning run: safety fine-tuning can erode as well as add, so it is tested after every change, not once.
Known limits
Like any language model it can produce confident but incorrect output. Check answers that matter.
Safety fine-tuning lowers the rate of harmful behavior; it does not reduce it to zero, and determined jailbreaks and prompt injection remain possible. Research has shown that even benign fine-tuning can erode safety behavior, which is why this model is evaluated after every change.
A safety-tuned model can be over-cautious and refuse reasonable requests. Over-refusal is measured and tracked alongside harmful compliance, not ignored.
Design intent is a statement of what the model is built to do. Only entries marked as measured, once published, describe what it has been shown to do.
Frequently asked questions
What is Swfte Safety?
Swfte Safety is the working name for Swfte’s own fine-tuned, safety-first language model, designed for governed agents, SecOps triage and regulated workflows. The base model, training data provenance, evaluation results, release date, licence and availability are still to be published.
Are there benchmark results?
No measured results are published yet. The design-intent table lists what will be measured, and each measured cell is a placeholder until a reproducible result exists. We do not publish numbers we cannot reproduce.
Is the model red-teamed?
Red-teaming is part of the release gate described on the open-source model testing page. Results are not published yet, so we make no claim about outcomes.
Does a safety-tuned model remove the need for guardrails?
No. Safety fine-tuning reduces harmful behavior but does not eliminate it. Runtime guardrails, access control and human approval remain part of the design.
Is it open-weight?
The licence and availability are not decided publicly yet. See the model card for the fields that remain to be published.
Does using it settle my regulatory position?
No model can. Swfte provides the technical controls, governance mechanisms and evidence you need to deploy AI within your applicable regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration.