# How to fine-tune an LLM on your own data

Canonical: https://www.swfte.com/how-to-fine-tune-an-llm-on-your-own-data
Last verified: 2026-10-06
Difficulty: Advanced
Time: About 1 to 2 weeks end to end, most of it on data and evaluation. The training run itself is usually the shortest part.
Cost: Dominated by people: the hours subject experts spend writing and reviewing examples. Compute is a smaller line. A formula is in step 7.
Hardware: None for the hosted route. For open weights: a GPU you rent or own with enough memory for your chosen model; see the sizing notes in the local model guide.

## Short answer

Fine-tune an LLM when a prompt cannot make it behave consistently on a narrow, stable task and you have reviewed examples of correct output. Define a held-out test first, build and clean a JSONL dataset, train either through a hosted fine-tuning API or with open weights using Hugging Face TRL, then compare against the untouched model. If the problem is missing knowledge, use retrieval instead.

## Who this is for

- Engineering and data leads deciding whether fine-tuning is worth doing for a specific task, and how to do it properly if it is.
- Teams with reviewed examples of correct output for a stable, high-volume task such as classification, extraction or fixed-format writing.
- Anyone who has read about fine-tuning and wants the dataset format, the training calls and the evaluation spelled out.

Not for:
- Teams whose problem is that the model lacks company knowledge. See [how to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system).
- Anyone who wants to run the whole thing on one laptop. [How to create your own local model](https://www.swfte.com/how-to-create-your-own-local-model) is the hands-on local version, with LoRA, GGUF and Ollama.

## Prerequisites

- A task defined in one sentence, with a written description of a correct output.
- At least a few hundred real input and output pairs that a qualified person has reviewed, and the right to train on them. Personal data removed or minimised.
- A baseline: how the current prompt performs on a fixed set of cases. Without it you cannot show an improvement.
- An owner for the finished model who will re-test and retire it.

## The decision rule: prompt, retrieve, then tune

Fine-tuning changes how a model behaves: its format, tone, label choices and consistency. It is a poor way to teach facts that change, because a fact trained into the weights goes stale and the model repeats it with full confidence. The shorthand most practitioners use is: retrieval for knowledge the model lacks, fine-tuning for behaviour it lacks, and a good prompt before either.

We cover the decision in depth in [when to fine-tune a model on your own data](https://www.swfte.com/blog/when-to-fine-tune-a-model-on-your-own-data-2026), and the safety side in [safety-first fine-tuning explained](https://www.swfte.com/blog/safety-first-fine-tuning-explained-2026). This guide assumes you have read the short version below and want the method.

| Symptom | First fix | Fine-tune when |
| --- | --- | --- |
| Wrong or outdated facts | Retrieval from a maintained source | Almost never: it is a knowledge problem |
| Inconsistent output format | Clear instructions, a schema and worked examples in the prompt | The format still drifts on a held-out set after a good prompt |
| Wrong label or classification on edge cases | Better examples in the prompt, then a rule-based check | You have hundreds of reviewed edge cases and the task is stable |
| Tone or house style | A style guide in the system prompt | A style guide plus examples is not consistent enough at your volume |
| Cost or latency of a large model on a narrow task | A smaller model with a good prompt | A small tuned model beats a small prompted one and the volume justifies owning it |

## Which training route?

There are two practical routes. The right one depends mostly on where your data may go and on how much control you need over the finished model.

|  | Hosted fine-tuning API | Open weights you train yourself |
| --- | --- | --- |
| Where data goes | To the provider for the training job; read their data terms first | Stays on hardware you choose |
| What you get back | A model id you can call only through that provider | Weights or an adapter you can keep, move and run anywhere |
| Effort | Low: upload a file, start a job | Higher: environment, training script, conversion, serving |
| Models available | Whatever the provider lists as tunable; this changes | Any open-weight model whose licence allows it |
| Exit plan | You depend on the provider for the tuned model | You own the artefacts; hosting is your choice |

> NOTE: Either way, keep your dataset and evaluation set in your own storage. A tuned model can be rebuilt on another route; a lost dataset cannot.

## Steps

### Step 1: Define the task, the test and the baseline

Outcome: A one-sentence task, a frozen held-out test set and a baseline score.

Write the task in one sentence with a subject, a verb and a boundary: classify inbound support tickets into five categories; extract invoice totals into a fixed JSON schema; draft a claims summary in the house format. Then write what a correct output looks like, in enough detail that two reviewers would agree.

Before any training, set aside a test set and do not touch it. Pick cases that include the hard and rare ones, not only the typical. A few dozen is enough to start; more makes the result steadier. Score the current prompt on it and record the number and the date. That number is the bar.

Choose the metric to match the output: exact match for labels and fixed formats, field-level accuracy for extraction, a rubric or human review for free text. If you plan to use a model as a judge, read the pitfalls in [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai) first. Decide in advance how much improvement counts as a win and what regressions disqualify the model.

> WARNING: Writing the pass rule after you have seen the numbers is how evaluations get bent to fit the model someone already likes. Write it first.

### Step 2: Build and clean the training dataset

Outcome: A cleaned JSONL training file and a separate validation file, with provenance recorded.

Training data is pairs of an input and the output you want. The common format is chat style: a list of messages ending with the assistant reply. Hugging Face TRL accepts conversational data like this, and also prompt and completion pairs, and applies the model chat template for you. OpenAI's supervised fine-tuning takes JSONL with a `messages` list in the same shape. One file can serve both routes.

Volume: OpenAI documents a minimum of 10 examples for its service and reports improvements from 50 to 100 examples, and advises starting with about 50 good ones. For open weights, plan on a few hundred to a few thousand reviewed examples as a working range; that is our guidance, not a documented threshold, and the held-out score is the real judge. More low-quality data is worse than less clean data.

Quality rules that matter more than volume. Cover the hard cases. Make labels consistent by settling disagreements in a written guideline. Remove duplicates and near-duplicates, which cause the model to over-weight one pattern. Remove personal data and anything you have no right to use. Record where each example came from, who approved it and when. Our page on [data preparation for custom models](https://www.swfte.com/platform/custom-models/data-preparation) sets out selection, sanitisation and consent.

Keep three sets apart from the start: train, validation (used to watch for overfitting during training) and the test set from step 1. If one customer, document or conversation yields many examples, split by that unit, or near-copies leak across the boundary.

One training example per line (chat format):

```json
{"messages": [{"role": "system", "content": "Classify the ticket as billing, bug or other. Reply with one word."}, {"role": "user", "content": "I was charged twice for March."}, {"role": "assistant", "content": "billing"}]}
```

Prompt and completion pairs also work with TRL:

```json
{"prompt": [{"role": "user", "content": "I was charged twice for March."}], "completion": [{"role": "assistant", "content": "billing"}]}
```

### Step 3: Route A: train through a hosted fine-tuning API

Outcome: A fine-tuned model id you can call through the provider's API.

Read the provider's data terms before uploading. OpenAI states that data sent to its API is not used to train or improve its models unless you opt in, and that abuse-monitoring logs are retained for up to 30 days by default. Those are statements about general API data; confirm how your training files are handled for the specific service and agreement you use, and repeat this check for any other provider.

OpenAI's fine-tuning guide lists several methods. Supervised fine-tuning, where you provide examples of correct responses, is what this guide uses; the others (vision fine-tuning, preference optimisation and reinforcement fine-tuning) exist for different needs. At the time of reading, the page named `gpt-4.1-2025-04-14`, `gpt-4.1-mini-2025-04-14` and `gpt-4.1-nano-2025-04-14` as supported for supervised fine-tuning. The tunable list changes, so check the current page before you plan around a model.

The flow is: upload the JSONL file with purpose `fine-tune`, create a job with the file id and the base model, wait for it to finish, then call the resulting model id like any other model. OpenAI documents the fine-tuned model id as a string beginning `ft:`. Evaluate it against your test set exactly as you would any other model, and keep its id and the training file id in your model record.

Upload the training file and start a job (OpenAI Python SDK):

```python
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from the environment

training_file = client.files.create(
    file=open("data/train.jsonl", "rb"),
    purpose="fine-tune",
)

job = client.fine_tuning.jobs.create(
    training_file=training_file.id,
    model="gpt-4.1-nano-2025-04-14",
)
print(job.id)
```

> NOTE: Model names and availability change. Take the model string from the provider's current fine-tuning page, not from this guide.

### Step 4: Route B: train open weights with Hugging Face TRL

Outcome: A LoRA adapter trained on your data that you own, ready to merge and serve.

TRL's `SFTTrainer` is the standard open-source way to run supervised fine-tuning. Pass it a model id and your dataset; add a `peft_config` to train a small LoRA adapter instead of every weight. The TRL docs suggest a higher learning rate, around 1e-4, when training adapters, and show that conversational datasets get the chat template applied automatically.

Choose a base model whose licence allows your use, because the tuned model inherits it. Read the licence on the exact checkpoint. Start small, in the 0.5B to 4B range, so a run takes minutes to hours. For memory sizing, QLoRA setup, merging the adapter and converting to GGUF, follow [how to create your own local model](https://www.swfte.com/how-to-create-your-own-local-model): that guide has the full commands and this one stays at the method level.

Watch validation loss. If it rises while training loss keeps falling, you are overfitting: use fewer epochs, more varied data or a lower rank. Training loss alone proves little. The only result that counts is the held-out comparison in the next step.

Install:

```bash
pip install trl peft datasets
```

Minimal adapter training (train.py):

```python
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer

data = load_dataset("json", data_files={"train": "data/train.jsonl", "valid": "data/valid.jsonl"})

trainer = SFTTrainer(
    model="Qwen/Qwen3-4B",
    train_dataset=data["train"],
    eval_dataset=data["valid"],
    peft_config=LoraConfig(r=16, lora_alpha=16, target_modules="all-linear"),
    args=SFTConfig(
        output_dir="runs/sft",
        learning_rate=1e-4,
        num_train_epochs=2,
        eval_strategy="steps",
        eval_steps=50,
    ),
)
trainer.train()
trainer.save_model("adapter-out")
```

### Step 5: Evaluate against the baseline, and against harm

Outcome: A written verdict: ship, iterate or stop, backed by held-out numbers.

Run the tuned model on the frozen test set from step 1, with the same prompts and settings as the baseline, and compare. Read the misses one by one: they tell you whether the data, the task definition or the base model is the limit. Compare on the slices that matter, such as the rare categories and the longest inputs, not only the average.

Then test what you did not train on. Fine-tuning can erode safety behaviour even when the new data is harmless, so run the same refusal, over-refusal and prompt-injection checks on both models and compare. Test every language you serve. Check general ability on a few tasks unrelated to your own. Make these checks gates: a candidate that improves the task score and regresses on safety does not ship. The method for building the suite, calibrating judges and writing the evidence pack is in [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai), and [evaluation and safety for custom models](https://www.swfte.com/platform/custom-models/evaluation-and-safety) covers the same ground in a governed deployment.

Record the result with the date, the model and dataset versions, the settings and who signed it off. If the improvement is smaller than the margin you set, do not ship it and call it a win. Improve the data, or go back to prompting and retrieval.

### Step 6: Work out what it really costs, then plan ownership

Outcome: A cost estimate that includes people and serving, and a named owner.

Compute is the visible cost and usually the smaller one. Estimate with a formula, using your own numbers. Total cost is roughly the labelling and review hours times their rate, plus training cost, plus the cost of running the model for as long as it stays in service, plus re-testing. Training cost on a hosted API is the number of training tokens times the provider's price per million training tokens, so read the price from their pricing page on the day. On rented GPUs it is GPU hours times the hourly rate.

A worked example with made-up numbers, to show the shape, not to quote: 1,000 examples of 600 tokens each is 600,000 tokens per pass; three passes is 1.8 million training tokens. If a provider charged a hypothetical 10 currency units per million training tokens, training would cost 18 units. If subject experts spent 60 hours reviewing at a hypothetical 50 units an hour, review costs 3,000 units. The review hours dominate the compute, which is the usual pattern.

Then the cost that people forget: running it. A model you host needs serving capacity and monitoring, and a model on a hosted API may be priced differently from the base model. If call volume on the task is low, the people cost of owning a custom model can exceed any saving. Name an owner who will re-test it when the base model, the task or the data changes, and who can retire it.

**Cost lines to estimate (use your own figures)**

| Line | How to estimate |
| --- | --- |
| Labelling and review | Hours x rate, for writing, reviewing and resolving disagreements |
| Evaluation harness | Engineering hours to build and maintain the test and safety suites |
| Training, hosted API | Training tokens x price per million training tokens, per run, including reruns |
| Training, open weights | GPU hours x hourly rate, per run, including reruns |
| Serving | Per-token price of the tuned model, or the cost of the server and its upkeep |
| Ownership | Re-testing, monitoring, retraining and retirement over the model's life |

### Step 7: Deploy behind a stable interface and monitor it

Outcome: The tuned model serving real traffic with a rollback path and a drift check.

Put the model behind an interface your application already uses so you can swap it. Many servers expose an OpenAI-compatible API, so the model name and base URL can live in configuration, which is what makes a rollback to the previous model a one-line change. For serving open weights, [how to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm) covers a production server, and a gateway ([how to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway)) gives you routing and fallback.

Roll out gradually: send a small share of traffic to the tuned model, compare against the old path on live cases, and keep the old path available. Keep the model record current: base model and revision, licence, dataset and test versions, settings, results and owner.

Set re-test triggers: a new base model version, a change to the task or policy, a drop in the live score, or a fixed review date. Fine-tuned models are not set and forget, because the facts and rules around them keep moving.

## When to stop

Stop before step 4 if the baseline already meets your threshold. Stop after step 6 if the tuned model does not beat the baseline by the margin you wrote down, or if it regresses on safety, refusals or any capability you rely on. Go back to retrieval if the misses are wrong facts. A failed run is a result: it tells you the task or the data is the problem, which is cheaper to learn before you build a deployment around the model.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| The tuned model is no better than the prompted base model | The data is too small or inconsistent, the task is really a knowledge problem, or the prompt baseline was weak. | Read the misses. Fix label disagreements, add hard cases, or move to retrieval. Re-run the baseline with a stronger prompt to make sure you are comparing fairly. |
| Excellent on the validation file, poor on real traffic | The validation set resembles the training set too closely, or leakage between them. | Rebuild the split by customer or document. Add a test set sampled from recent live inputs. |
| Validation loss rises while training loss falls | Overfitting. | Use fewer epochs, more varied examples or a smaller adapter rank. Stop at the best validation point. |
| The model confidently states outdated facts | Facts were trained into the weights and then changed. | Move changing knowledge to retrieval. Keep fine-tuning for behaviour. See [how to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system). |
| The model got worse at refusing unsafe requests, or over-refuses | Fine-tuning shifts safety behaviour even on harmless data. | Run refusal and over-refusal suites on base and tuned models side by side, gate the release, and mix in general-purpose examples or reduce training strength. |
| The hosted job fails at validation of the training file | A malformed JSONL line, an empty message or a wrong role. | Check each line parses as JSON and ends with an assistant message. Read the provider's error for the line number. |
| Output never terminates or repeats after local conversion | Chat template or end-of-turn token mismatch between training and serving. | Use the same template when you serve as when you trained, and test the unconverted model first. See the local model guide. |

## Verify it worked

- [ ] The task, the pass rule and the baseline score were written down before training.
- [ ] Train, validation and test sets share no examples and the test set was never used to tune.
- [ ] The tuned model beats the baseline on the held-out test by the margin you set.
- [ ] Safety, refusal and over-refusal checks pass on the tuned model, in every language you serve.
- [ ] Every example has recorded provenance and approval, and personal data was removed or minimised.
- [ ] A named owner, a rollback path and re-test triggers exist.

## Next steps

- [How to create your own local model](https://www.swfte.com/how-to-create-your-own-local-model): the hands-on LoRA, GGUF and Ollama path on your own hardware
- [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai): build the evaluation suite, regression gates and evidence pack
- [How to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system): the right fix when the model lacks knowledge
- [When to fine-tune a model on your own data](https://www.swfte.com/blog/when-to-fine-tune-a-model-on-your-own-data-2026): the decision ladder and readiness signals in more depth

## FAQ

### How many examples do I need to fine-tune an LLM?

It depends on the task. OpenAI documents a minimum of 10 examples for its hosted service and reports gains from 50 to 100. For open weights, a few hundred to a few thousand reviewed examples is a reasonable working range. Judge by the held-out score, and favour clean, varied data over volume.

### Should I use RAG or fine-tuning?

Use RAG when the model lacks knowledge that exists in your documents and changes. Use fine-tuning when it has the information but behaves inconsistently on a narrow, stable task. Many systems use a good prompt, retrieval for facts and a tuned model only for the behaviour retrieval cannot fix.

### What format does fine-tuning data need to be in?

Usually JSONL with one example per line. The common shape is a messages list ending with the assistant reply, which OpenAI and Hugging Face TRL both accept. TRL also takes prompt and completion pairs, and applies the model chat template to conversational data.

### Is fine-tuning safe for my data?

It depends on the route. On open weights the data stays on hardware you choose. On a hosted API it goes to the provider, so read their data terms; OpenAI states API data is not used for training unless you opt in. Remove personal data you do not need before any route.

### Can fine-tuning make a model less safe?

Yes. Further training can shift refusal and safety behaviour even when the new data is harmless. Run the same safety and refusal suite on the base and the tuned model, compare them side by side and treat a regression as a failed run.

### How much does fine-tuning cost?

Estimate labelling hours, evaluation work, training and serving, plus ongoing re-testing. People time usually exceeds compute. Hosted training is training tokens times the provider price per million tokens; rented GPUs are hours times the hourly rate. Step 7 has a worked example with made-up numbers.

### Does Swfte offer managed fine-tuning?

No, not yet. Swfte has a Model Vault designed to hold and deploy weights you bring, but no training service. The training routes in this guide use a hosted provider API or open-source tools you run yourself.

## How Swfte can help

You can follow this guide with no Swfte account. Swfte can help around the edges: keeping and deploying weights you produce, and routing between a tuned model and others behind one API.

- [Custom models](https://www.swfte.com/platform/custom-models): how Swfte approaches models an organisation owns
- [Domain fine-tuning](https://www.swfte.com/platform/custom-models/domain-fine-tuning): the methods compared in a governed deployment
- [Data preparation](https://www.swfte.com/platform/custom-models/data-preparation): selection, sanitisation and consent for training data
- [Connect](https://www.swfte.com/products/connect): an OpenAI-compatible gateway to route to a tuned model and fall back to another

Swfte does not offer managed fine-tuning today, and there is no Swfte training service to buy. Where this site discusses hosting your own models, availability and plans are <custom model availability - founder to fill>. If another page on this site describes a fine-tuning flow inside Swfte, treat this guide as the current statement.

## Sources

- [TRL SFTTrainer documentation](https://huggingface.co/docs/trl/sft_trainer): dataset formats (conversational and prompt-completion), SFTTrainer, peft_config, adapter learning rate near 1e-4
- [PEFT LoRA developer guide](https://huggingface.co/docs/peft/main/en/developer_guides/lora): LoraConfig fields r, lora_alpha and target_modules
- [OpenAI model optimization guide](https://developers.openai.com/api/docs/guides/model-optimization): fine-tuning methods and the models listed for each, as read on 2026-10-06
- [OpenAI supervised fine-tuning guide](https://developers.openai.com/api/docs/guides/supervised-fine-tuning): JSONL messages format, 10-example minimum, gains from 50 to 100 examples, files.create and fine_tuning.jobs.create calls, ft: model id form
- [OpenAI data controls](https://developers.openai.com/api/docs/guides/your-data): API data not used for training unless opted in; abuse-monitoring logs up to 30 days
- [Swfte: when to fine-tune a model on your own data](https://www.swfte.com/blog/when-to-fine-tune-a-model-on-your-own-data-2026): decision ladder and readiness signals (our own post)

Last verified against these sources on 2026-10-06.
