← The journal
Technology

When Should You Fine-Tune a Model on Your Own Data?

When to fine-tune a model on your own data: to change behaviour and format, not to teach facts. A decision ladder.

Swfte Journal / Technology

Fine-tune a model on your own data when you need to change how it behaves: the format it writes in, the tone it holds, the narrow decision it has to make the same way every time. Do not fine-tune to teach it facts that change, such as prices, policies, account details or this quarter's org chart. Facts belong in a retrieval layer you can update in an afternoon. Weights are slow to change, expensive to re-test and stubborn about forgetting.

That one distinction settles most of the arguments teams have about fine-tuning. The rest of this post is about the cases it does not settle: how to tell whether you are ready, what data you need, which method fits, how safety can quietly slip, and what owning the result involves. For the wider picture of how domain models fit into a governed platform, start with custom models on Swfte.

What does fine-tuning actually change?

A pretrained model has absorbed a great deal of general language. Fine-tuning nudges its weights using examples of the work you want done, so that its default response moves closer to your examples. It is very good at shifting habits. It is poor at installing reliable knowledge.

Consider a claims team that wants every assessment summary to follow the same seven headings, name the policy clause it relied on, and never speculate about fault. A well-written prompt gets most of the way there. Over thousands of cases, the model still drifts: a heading renamed, a clause omitted, a sentence of speculation slipped in. That drift is a behaviour problem, and fine-tuning is a reasonable fix.

Now suppose the same team wants the model to know the current excess for every product line. That is a knowledge problem. Train it in, and the model will state last year's excess with full confidence after the figure changes. Retrieval, as covered in our RAG architecture guide, handles this properly because the source of truth stays outside the weights.

Which rung of the ladder are you on?

Treat the options as a ladder. Climb one rung at a time, and only move up when the rung below has a clear, written reason for failing.

RungUse it whenThe test before moving up
Better promptThe behaviour can be described in wordsHave you tried clear instructions, a few worked examples and an output schema, and measured the result on a fixed set?
RetrievalThe model lacks information that exists somewhere and changesDoes the right passage reach the model, and does the answer still go wrong when it does?
Fine-tuneThe model has the information but behaves inconsistently on a stable, narrow taskCan you show the failure on a held-out set, with examples of correct output for the same inputs?
Train from scratchYou need a model for a modality or language no open base coversHave you priced the data, compute and research staff, and confirmed no adaptable base exists?

The last rung is rare. Almost every organisation that thinks it needs its own foundation model needs a better retrieval index and a small adapter. Training from scratch makes sense for a handful of research labs and very large institutions with unusual data. For everyone else, the useful question is whether to move from the second rung to the third.

The earlier post on fine-tuning an LLM on your own data walks through the general prompt, retrieve, tune method in more depth if you want a second explanation of the same ladder.

What are the signs you are ready to fine-tune?

Four conditions should all be true before anyone books GPU time.

  • The task is stable. The definition of a good answer has not changed in months and is not about to. If the policy behind the task is under review, wait.
  • You have clean examples. Real inputs paired with outputs a qualified person would sign off. Not synthetic guesses, not last year's outputs nobody checked.
  • You have an evaluation set. A held-out collection of cases, never used for training, with a scoring method agreed in advance. Without it you cannot tell whether the tuned model is better or only different.
  • Someone owns it. A named person or team responsible for the model after launch: re-testing it, retiring it, answering for its mistakes. A model with no owner becomes a liability the day the person who trained it moves on.

If you can tick all four, fine-tuning is worth a scoped trial. If you cannot, the work is upstream of training.

What are the signs you are not ready?

The warning signs are usually visible early. People disagree about what a correct output looks like. The examples come from a single enthusiastic reviewer. Nobody has measured how the current prompt performs, so there is no baseline to beat. The real complaint is that the model "does not know" something, which points back to retrieval. Or the task touches a decision that is still being argued about in a policy meeting.

Any of these means a tuned model would freeze an unsettled question into weights. That makes the eventual correction harder, not easier.

What data do you need?

The honest answer is qualitative, because the right amount depends on the task and on how far the target behaviour sits from what the base model already does.

What matters more than volume:

  • Coverage of the hard cases. Easy examples teach little. The examples near a decision boundary, the awkward edge cases and the inputs that currently fail carry most of the signal.
  • Consistency. If two reviewers would label the same input differently, the model learns the disagreement. Resolve it in a written guideline first.
  • Provenance. Know where each example came from, who approved it and whether you have the right to train on it. Personal data should be removed or minimised before it goes anywhere near a training run.
  • Separation. Keep the evaluation set apart from the training set from day one. Leakage between them makes every later result look better than it is.

Our page on data preparation for custom models sets out how selection, sanitisation and consent fit together, and which parts of that are built today.

LoRA, full fine-tune or distillation: what is the difference?

Three methods cover most practical work. In plain words:

LoRA (low-rank adaptation) freezes the original weights and trains a small set of extra parameters alongside them. The result is a compact adapter file that sits on top of the base. It is cheaper to train, easier to roll back and less likely to wipe out what the base model already did well. For format, style and narrow classification it is usually the first thing to try.

Full fine-tuning updates every weight in the model. It can move further from the base, which helps for a deep shift in vocabulary or reasoning, and it costs more in compute, storage and testing. It also has more room to damage general ability and safety behaviour. Reach for it when an adapter has clearly stopped improving.

Distillation trains a smaller model to imitate a larger one on your task. You run the larger model over your inputs, keep the outputs you accept, and train the smaller model on them. It suits cases where you have plenty of inputs but few labelled answers, and where you want a cheaper, faster model for a narrow job. Check that the larger model's terms allow its outputs to be used for training.

The domain fine-tuning page compares these methods in the context of a governed deployment.

How can fine-tuning erode safety, and how do you test for it?

This is the part teams most often skip. Safety behaviour in an aligned model, such as declining harmful requests or asking before a risky action, is learned during post-training. Further training can move it, even when the new data contains nothing harmful. A model tuned to be terse and decisive on claims summaries may also become more willing to answer questions it previously declined.

So a tuned model is a new model, and it needs its own safety evaluation. In practice:

  1. Run the same safety and refusal suite on the base and on the tuned candidate, and compare them side by side.
  2. Measure over-refusal as well. A model that declines legitimate domain questions is a failure too.
  3. Test prompt injection through the channels your deployment uses: documents, tool output, retrieved pages.
  4. Repeat in every language you serve, since safety training is often thinner outside English.
  5. Make the comparison a gate. If the candidate regresses on safety, it does not ship, however good its task score.

We cover the research behind this in safety-first fine-tuning explained, and the evaluation approach for custom models on evaluation and safety.

Whose licence applies to your tuned model?

Every open-weight checkpoint ships with its own licence, and licences differ between model families and sometimes between sizes or versions within a family. Your tuned model inherits the terms of the exact base you started from. Read the licence file that ships with those specific weights, not a summary on a model card or a blog post.

Some licences carry use restrictions or attribution duties; some restrict certain regions or certain multimodal variants. Permissive licences such as Apache 2.0 and MIT are generally the simpler starting point for commercial work. That is a recommendation, and your legal team should confirm it for your case.

What does fine-tuning really cost?

Compute is the visible cost and usually the smaller one. An adapter run on a modest open model is not expensive. The larger costs are people: subject experts writing and reviewing examples, someone resolving labelling disagreements, an engineer building the evaluation harness, a reviewer signing off each candidate, and the owner who keeps re-testing the model for as long as it is in service.

Then there is the cost of running it. A model you host needs serving capacity, monitoring and a plan for retirement. If the volume on the task is low, the people cost of a custom model can outweigh any saving on inference. Fine-tuning pays off on stable, high-volume, narrow tasks where consistency matters, and rarely elsewhere.

What can you do on Swfte today?

Here is the honest split. Hosting and serving your own weights is built. Model Vault lets you upload weights you have trained elsewhere, records a checksum manifest and the base model they came from, keeps versions, promotes them through development, staging and production, and deploys them to a dedicated endpoint with an audit log you can export. Connect can then sit in front of that endpoint, so applications reach it through one OpenAI-compatible API with routing, usage caps, cost tracking and audit events, alongside any frontier model you also use. Single-turn evaluation in Studio lets you score a candidate on chat-style cases.

Managed adaptation runs, where Swfte would train the adapter or full fine-tune for you, are on the roadmap, as are data export from the company brain, consent capture, named domain evaluation suites and automated red-teaming. The supported base models and adaptation methods for that service are not published yet. Today, you bring the trained weights; Swfte helps you host, version, route and audit them.

If you are weighing whether to start, write down the task, the four readiness signals and your evaluation set first. If those come together easily, a small adapter trial is cheap to run and easy to discard. If they do not, the work you need is in the data and the definition of the task, and no amount of training will substitute for it.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Centralise your knowledge in Cortex

The desktop AI workspace: 20+ providers, local models, knowledge bases with RAG that cite their sources, MCP tools and agents, with sensitive work staying on the laptop by default.