Safety-First Fine-Tuning: What It Is, What It Cannot Do
How safety-first fine-tuning works, why fine-tuning can erode safety and how to test refusal and over-refusal.
Most language models are tuned to be helpful first and careful second. That ordering is fine for a chat window with a human reading every reply. It is the wrong ordering for an agent that reads a ticket, calls three tools and writes to a system of record, because the model will follow instructions from anyone or anything that can put text in front of it. Safety-first fine-tuning reverses the priority: it treats declining, asking and escalating as first-class behaviors to be trained and measured, not as filters bolted on afterwards.
This post explains what that means in practice, what the research says about how fragile safety behavior is, how to evaluate it honestly, and where it fits in a layered design. It is the background for Swfte Safety, the safety-first model we are designing, and it is deliberately specific about what is a design intent and what has been measured. For Swfte Safety, no measured results are published yet.
What "safety-first fine-tuning" actually means
A pretrained base model predicts text. It has no particular tendency to refuse anything. Almost every deployable model is made by post-training: supervised fine-tuning on example conversations, then preference-based tuning such as reinforcement learning from human feedback or direct preference optimization, which pushes the model toward responses people (or other models) rated higher. Safety behavior is one of the things this stage teaches.
"Safety-first" is a statement about what the training objective rewards and in what order. In a safety-first recipe, three kinds of behavior are trained deliberately rather than left to emerge:
- Refusal that is calibrated. Decline clearly harmful requests, briefly, without a lecture, and offer a safe alternative when one exists.
- Escalation. When an action is out of policy, ambiguous or high-impact, ask a clarifying question or hand off to a human instead of guessing. For agents this matters more than refusal, because the useful failure mode is "I need approval", not "no".
- Policy-following. Prefer a system-level policy (allowed actions, restricted actions, approval rules) over conflicting instructions that arrive in the user turn or inside retrieved content.
None of this makes a model safe in an absolute sense. It shifts probabilities. The point of a safety-first model is that the default, when things get ambiguous, is the cautious behavior rather than the compliant one.
Safety behavior is easier to erode than to add
The most important research result for anyone planning to fine-tune is that safety alignment is shallow in a specific, measurable way. In "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" (Qi et al., first posted to arXiv in October 2023), the authors showed that fine-tuning GPT-3.5 Turbo on as few as 10 adversarially designed examples, at a cost of less than $0.20 through the provider's API, was enough to jailbreak it. They also found that fine-tuning on benign, commonly used datasets inadvertently degraded safety alignment, to a lesser extent.
Read that as a design constraint. Three consequences follow:
- Fine-tuning is not a neutral operation. Any customization of an aligned model, including a well-meant one on clean business data, can move refusal behavior. You have to re-measure afterwards.
- A clean dataset is not enough. A training set with no harmful examples can still shift the model. Evaluate the model, not just the data.
- Every later change inherits the risk. A new quantization, a new base revision or another fine-tuning run is a new candidate, not a patch.
This is why a safety-first pipeline ends with an evaluation gate rather than with a training run. Our open-source model testing methodology applies the same gate to every model we would deploy, including our own.
How to evaluate a safety-tuned model without fooling yourself
A single "safety score" hides the trade-offs. Measure at least these, separately:
| Property | The question | How to probe it |
|---|---|---|
| Harmful-request refusal | Does it decline clearly harmful requests? | Curated harmful-behavior sets such as HarmBench, which reports attack success rate |
| Over-refusal | Does it wrongly decline benign requests that sound risky? | XSTest-style pairs: safe prompts that resemble unsafe ones, next to genuinely unsafe contrasts |
| Jailbreak resistance | Does refusal survive role-play, encodings and multi-turn pressure? | Automated attack tools and manual red-teaming |
| Prompt-injection resilience | Does it treat retrieved text as data? | Instructions hidden in documents, tool output and web pages |
| Policy-following | Does a system policy beat a conflicting user instruction? | Conflict cases written from your own policy set |
| Language parity | Does behavior hold outside English? | The same suites in each language you serve |
Two of these pull against each other. A model that refuses everything scores perfectly on harmful-request refusal and is useless; a model that never refuses is helpful and dangerous. XSTest was designed for exactly this tension: its safe prompts contain sensitive-sounding words, such as asking how to "kill a process", and are paired with genuinely unsafe prompts, so you can see whether the model understands intent rather than reacting to keywords. A published caveat applies, though: older over-refusal sets get easier for new models, so keep a private set that reflects your own domain.
Report sample sizes next to every score, record the harness version, prompt template and sampling settings, and keep a held-out set that you never train on. A clean run on a known attack corpus means those attacks did not succeed. It does not mean the model is safe.
Why language coverage is part of safety
Safety training data is usually richest in English. In practice that means refusal and injection resilience can be weaker in other languages, and a model that behaves well in English can be easier to push around in Polish or Portuguese. For an EU deployment, this is a safety question as much as a quality question. Run the same suites in the EU languages you serve, compare each against English, and treat a gap as a defect to fix or a limit to disclose. We cover the deployment side in deploying an LLM in the EU.
Where fine-tuning sits in the EU's rules
The EU AI Act's obligations for general-purpose AI model providers apply from 2 August 2025. The European Commission's guidelines give an indicative test for when someone who modifies a model becomes a provider in their own right: the compute used for the modification exceeds one third of the compute used to train the original model. A light fine-tune is unlikely to cross that line, but a large continued-training run might. Whether your system is high-risk depends on its use, not on the model inside it, and the timing of those rules has been shifting, so check the current text. None of this is legal advice. The practical point is simple: keep a record of what you changed, from which base, and what you measured afterwards.
Fine-tuning is one layer, not the defense
Even a well-evaluated safety-tuned model is one layer. The others do different jobs:
- Runtime guardrails check inputs and outputs and limit what tools a model-driven agent can call. They answer "should this be blocked?"
- Governance answers who was acting, allowed to do what, under which policy, with which data, using which model, and whether you can prove it. In a Swfte deployment that is the job of Nexus and the policy set around it.
- Human approval gates consequential actions, so a model that is talked into something still cannot do it alone.
- The evaluation gate stops a model from being promoted at all if it regresses.
Prompt injection is the reason you cannot rely on the model alone. It is ranked first in the OWASP Top 10 for LLM Applications 2025, and the foundational research on indirect injection (Greshake and colleagues, 2023) found that mitigations were lacking when the paper was published. A safety-tuned model that treats retrieved content as data helps, and so do the other layers. See our AI governance overview and MCP security best practices for the surrounding controls.
What we are and are not claiming about Swfte Safety
We would rather be precise than impressive. Swfte Safety is a working name for a model designed around the behaviors above, for governed agents, SecOps triage and regulated workflows. Its base model, training data provenance, evaluation results, release date, licence and availability are not published, and the model card shows them as open fields rather than guesses. The design-intent table says what the model is meant to do; the measured column stays empty until there is a reproducible result to put in it.
If you are planning your own safety-tuned model, the checklist is short. Fix the behaviors you want in writing. Build the evaluation set before the training set. Measure refusal and over-refusal together. Re-run everything after every change. Keep guardrails, governance and human approval in place regardless. And when you deploy, follow the same loop for every model: Pick, Evaluate, Harden, Deploy, Monitor.
Related reading: how to evaluate an open-source LLM before production, the LLM red-teaming checklist, and the broader AI safety report guide for enterprises.