Platform / Custom models / Domain fine-tuning
Domain fine-tuning: when to adapt an open-weight model, and how to do it safely
A plain guide to choosing between prompting, retrieval, fine-tuning and training from scratch, and to adapting an open-weight base on your own data without losing safety along the way.
Most teams that ask for a custom model need something smaller: a better prompt, or retrieval over documents that change every week. Some genuinely need a model that behaves differently, writes in their format and handles a narrow task reliably. This page sets out the four levers, a rule for choosing between them, the adaptation methods in plain words, how to pick an open-weight base, the risks that come with any run, and what Swfte offers today against what it is designed for.
The four levers, and when each earns its cost
Each lever costs more than the one before it, in effort, in evidence you must keep, and in what you have to re-test when something changes.
Prompting
Instructions, examples and a fixed system policy. Cheapest to change and easiest to review. Start here, and measure it properly before deciding it is not enough.
Retrieval
The model reads relevant passages at request time. Right for facts that change, for answers that need a citation, and for content that different people are allowed to see differently.
Fine-tuning
The model's weights, or a small adapter on top of them, are trained on your examples. Right for behaviour, style, output format and reliability on a narrow, repeated task.
Training from scratch
A new model built from raw data. Rarely justified. It makes sense when the deliverable is the licence and the full data lineage themselves, and the budget and data to match exist.
A decision rule you can apply in a meeting
Fine-tune for behaviour, retrieve for facts. If the problem is that the model answers in the wrong shape, uses the wrong register, ignores your house format or wobbles on a narrow classification it sees thousands of times a day, adaptation is the right tool. If the problem is that the model does not know this quarter's policy, a new product or a customer's contract, retrieval is the right tool, because those facts will change next month and a trained model would quietly keep the old version.
Most real systems use both. A model adapted to your domain vocabulary and output format still reads current documents at request time, and still answers only from what the asking person is allowed to see. Training from scratch sits apart from the other three. Choose it only when the thing you need to own is a model whose every training document you can account for and whose licence is yours alone, and accept that this is a research programme, not a project.
One practical test: write down ten failures of the current system. If most are wrong facts, build retrieval. If most are wrong behaviour on correct facts, consider adaptation. If you cannot write down ten, you are not ready to train anything.
Adaptation methods in plain words
The method decides what you end up owning, what hardware you need and how much of the base model you can disturb.
LoRA adapters
Small trainable matrices added beside the frozen base weights. The adapter is a separate, small file. Cheap to train, easy to swap, and the base model stays exactly as published.
QLoRA
The same idea with the frozen base loaded in a compressed, quantised form during training, so a larger base fits on less memory. Quality must be checked on the artifact you will actually serve.
Full fine-tune
Every weight is updated. More capacity to change behaviour, more hardware, and more risk of disturbing what the base already did well, including its safety behaviour.
Continued pretraining
Further training on a large body of unlabelled domain text, so the model absorbs vocabulary and phrasing. Usually followed by instruction tuning, and usually the most expensive option short of a new model.
Distillation
A smaller model is trained on the outputs of a larger one, to keep most of the useful behaviour at lower serving cost. Check that the larger model's terms allow its outputs to be used this way.
Picking an open-weight base
Licence comes first, and it is per checkpoint, not per family. Read the licence file that ships with the exact weights you download and record the revision. As a recommendation, start from bases released under permissive licences such as Apache 2.0 or MIT, because derivative weights stay easy to deploy and to sell on. The Llama licence carries an EU restriction on some multimodal models, set out on the deploy-models hub; have counsel read the current text before an EU deployment.
Then size against hardware. A base that only fits your serving hardware at aggressive quantisation is a base you will spend months fighting. Check language coverage for every language your users write in, English included, and test it rather than trusting a model card. Finally, consider EU suitability as a whole: where the weights came from, whether you can host them in region, and whether the maker's terms let you keep everything inside your own infrastructure. Families such as Qwen, Mistral, Gemma and DeepSeek are common starting points; which bases Swfte supports for customer adaptation is <supported base models - founder to fill>.
The risks that come with every run
Overfitting
The model memorises the training set and does worse on anything slightly different. Held-out evaluation cases, drawn from the same task but never trained on, are the only reliable check.
Forgetting
Narrow training can erode general ability the base had, such as reasoning, other languages or following unusual instructions. Re-test the general capabilities you rely on as well as the target task.
Safety erosion
Even fine-tuning on benign data can weaken refusal behaviour that the base model was trained to have. Every adapted model is a new candidate and goes back through the safety suite.
Licence contagion
Derivative weights inherit obligations from the base licence, and sometimes from the licence of any model whose outputs were used as training data. Record both, before the run.
What to put in a brief before any run
If these answers are not written down, the run is not ready to start.
- 01
The task and its failure list
What the model must do, the failures of the current system that justify training, and why prompting and retrieval were not enough.
- 02
The base and its licence
The exact checkpoint and revision, the licence file read and recorded, and who signed off on its terms.
- 03
The data and its lawful basis
Which sources, selected by purpose and by access, what was removed during sanitisation, and the organisation's own decision on lawful basis.
- 04
The evaluation plan
Domain cases, held-out cases, safety and multilingual suites, and the result the current production model achieves on all of them.
- 05
The promotion rule
What result lets the candidate move from development to staging to production, and who approves each step.
- 06
The exit and retirement plan
How the model is rolled back, how it is retrained when data is erased, and how it is retired.
A worked example: a regulated team and its own vocabulary
Consider a compliance operations team in a regulated firm. Every day it receives incoming correspondence that must be sorted into a fixed set of categories and answered with a first draft in the firm's own terminology and structure. A general model, well prompted, gets the categories mostly right but keeps drifting into generic phrasing and occasionally invents a category that does not exist.
The team writes down its failures and finds they are behaviour, not facts: the model knows enough, but does not hold the format or the vocabulary. So it chooses a small open-weight base under a permissive licence, records the revision, and assembles labelled examples from past correspondence its reviewers had already checked, after sanitisation and with its legal team's decision on lawful basis on file. It trains a LoRA adapter, keeps a held-out set the adapter never saw, and runs the domain, safety and multilingual suites against both the adapter and the prompted model in production.
The adapter wins on the domain cases and holds its refusal behaviour, so it is promoted to staging, then to production, with the old configuration kept routable. Policy references still come from retrieval, because they change. The draft still goes to a person before it is sent.
What Swfte does today, and what it is designed for
Today Swfte hosts and serves models you have adapted yourself. The Model Vault takes your weights, records a hash manifest and the base model, versions them and promotes them through development, staging and production to a dedicated endpoint, and Connect puts the governed API in front. Swfte does not run fine-tuning jobs for customers today. A small proof-of-concept adapter pipeline has been run internally, at small scale; it is not a service.
Managed adaptation runs, with data selected from the company brain, evaluated before promotion and deployed into the vault, are designed for and on the roadmap. Adaptation methods that would be offered: <supported adaptation methods - founder to fill>. Availability and plans: <custom model availability and plans - founder to fill>.
Domain fine-tuning: built, in progress and roadmap
| Capability | Status | Notes |
|---|---|---|
| Host your own adapted weights in the Model Vault | Built | Upload, hash manifest, base model recorded, versions and stage promotion. |
| Serve behind the Connect gateway | Built | One OpenAI-compatible API with routing, usage caps, cost tracking and audit events. |
| Single-turn evaluation in Studio | Built | Scores single-turn chat against your cases. |
| Pre-model sanitisation gateway | In progress | Removes material that must not reach a model. Not yet something to rely on. |
| Managed LoRA or full fine-tune runs | Roadmap | Designed for. Swfte does not run training jobs for customers today. |
| Training data selected from the company brain | Roadmap | The brain does not export training data today. |
| Adapter serving as adapters | Roadmap | Whether adapters are served separately from merged weights is not yet settled. |
| Training in EU regions | Roadmap | Regions where training would run are not yet published. |
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Where this fits in the loop
Fine-tuning is the custom model station of the loop: an open-weight base adapted on data chosen with care, ready to be evaluated before it serves anyone.
- 01Company brainHolds what the organisation knows, with evidence statuses, history and access rules.
- 02Custom modelAdapted on data chosen from the brain, then evaluated and hardened before it ships.(this page)
- 03Governed agentsUse the model and read the brain, inside a Trust Profile, with approval where it matters.
- 04OutcomesWhat happened: approvals, corrections, results and cost, all on the record.
The four arrows
- Company brain to Custom model: select, sanitise, adaptRoadmap
Choose training data from the brain, remove what must not reach a model, adapt an open-weight base. The sanitisation gateway is in progress, and the data selection and training steps are on the roadmap.
- Custom model to Governed agents: serve, governBuilt
Serve the model on dedicated infrastructure behind the Connect gateway and bring agents onto it under policy. Model hosting and the gateway are built.
- Governed agents to Outcomes: act, recordBuilt
Agents act within their Trust Profile, with human approval for consequential steps, and every action is recorded.
- Outcomes to Company brain: written back as evidenceRoadmap
Outcomes return to the brain as new evidence with a status, and they decide when the model needs retraining. The write-back is on the roadmap.
Legend
- Built. Exists today and can be used.
- In progress. Being built. Not yet something to rely on.
- Roadmap. Designed for and on the roadmap. Not built. No dates are given.
Frequently asked questions
Should we fine-tune or use retrieval?
Fine-tune for behaviour, style, output format and reliability on a narrow repeated task. Use retrieval for facts that change, for answers that need a source, and for content with different access rules. Most production systems end up using both, with the adapted model reading current documents at request time.
What is the difference between LoRA and a full fine-tune?
LoRA trains a small adapter beside frozen base weights, so it is cheap and easy to swap, and the base stays untouched. A full fine-tune updates every weight, which gives more room to change behaviour but needs more hardware and carries more risk of disturbing what the base already did well.
Does Swfte run fine-tuning jobs for us today?
No. Swfte hosts and serves weights you have adapted yourself, through the Model Vault and Connect. Managed adaptation runs are designed for and on the roadmap. A small internal proof-of-concept adapter pipeline exists, but it is not a service.
Which base model should we start from?
Start from the licence of the exact checkpoint, then hardware fit, language coverage and whether you can host it in region. As a recommendation, prefer bases released under permissive licences such as Apache 2.0 or MIT. The bases Swfte supports for customer adaptation are not yet published.
Can fine-tuning make a model less safe?
Yes. Published research has found that refusal behaviour can weaken even after fine-tuning on harmless data. That is why every adapted model is treated as a new candidate and goes back through the safety suite, including jailbreak, harmful content and over-refusal checks, before promotion.
Is a fine-tuned model hosted on Swfte compliant with GDPR or the EU AI Act?
Swfte does not claim that any model or deployment is compliant. It provides technical controls and evidence that help an organisation meet its own obligations, such as recorded base models, versioned weights and audit logs. The posture depends on use case, jurisdiction, deployment and configuration, and on decisions your legal team makes.
Do we own an adapter we train on a licensed base?
You own your adapter weights, your training data and your evaluation sets. The base model stays under its maker's licence, and that licence still applies to anything derived from it. Record both before the run, and have counsel read the base licence for your use.
Take domain fine-tuning further with Swfte
Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.