Build · Advanced

How to fine-tune an LLM on your own data

  • Time: About 1 to 2 weeks end to end, most of it on data and evaluation. The training run itself is usually the shortest part.
  • Cost: Dominated by people: the hours subject experts spend writing and reviewing examples. Compute is a smaller line. A formula is in step 7.
  • Level: Advanced
On this page
  1. Short answer
  2. Before you start
  3. The decision rule: prompt, retrieve, then tune
  4. Which training route?
  5. 1. Define the task, the test and the baseline
  6. 2. Build and clean the training dataset
  7. 3. Route A: train through a hosted fine-tuning API
  8. 4. Route B: train open weights with Hugging Face TRL
  9. 5. Evaluate against the baseline, and against harm
  10. 6. Work out what it really costs, then plan ownership
  11. 7. Deploy behind a stable interface and monitor it
  12. When to stop
  13. Troubleshooting
  14. Verify it worked
  15. Next steps
  16. FAQ
  17. How Swfte can help
  18. Sources and last verified

Short answer

Fine-tune an LLM when a prompt cannot make it behave consistently on a narrow, stable task and you have reviewed examples of correct output. Define a held-out test first, build and clean a JSONL dataset, train either through a hosted fine-tuning API or with open weights using Hugging Face TRL, then compare against the untouched model. If the problem is missing knowledge, use retrieval instead.

The steps at a glance

  1. Define the task, the test and the baseline
  2. Build and clean the training dataset
  3. Route A: train through a hosted fine-tuning API
  4. Route B: train open weights with Hugging Face TRL
  5. Evaluate against the baseline, and against harm
  6. Work out what it really costs, then plan ownership
  7. Deploy behind a stable interface and monitor it

Before you start

Who this is for

  • Engineering and data leads deciding whether fine-tuning is worth doing for a specific task, and how to do it properly if it is.
  • Teams with reviewed examples of correct output for a stable, high-volume task such as classification, extraction or fixed-format writing.
  • Anyone who has read about fine-tuning and wants the dataset format, the training calls and the evaluation spelled out.

Probably not for you if

Prerequisites

  • A task defined in one sentence, with a written description of a correct output.
  • At least a few hundred real input and output pairs that a qualified person has reviewed, and the right to train on them. Personal data removed or minimised.
  • A baseline: how the current prompt performs on a fixed set of cases. Without it you cannot show an improvement.
  • An owner for the finished model who will re-test and retire it.
Time
About 1 to 2 weeks end to end, most of it on data and evaluation. The training run itself is usually the shortest part.
Cost
Dominated by people: the hours subject experts spend writing and reviewing examples. Compute is a smaller line. A formula is in step 7.
Hardware
None for the hosted route. For open weights: a GPU you rent or own with enough memory for your chosen model; see the sizing notes in the local model guide.
Skill
Comfortable with Python and JSON. Machine learning background helps but is not required.

Estimates are ours, not measurements, and move with your hardware, data and network.

The decision rule: prompt, retrieve, then tune

Fine-tuning changes how a model behaves: its format, tone, label choices and consistency. It is a poor way to teach facts that change, because a fact trained into the weights goes stale and the model repeats it with full confidence. The shorthand most practitioners use is: retrieval for knowledge the model lacks, fine-tuning for behaviour it lacks, and a good prompt before either.

We cover the decision in depth in when to fine-tune a model on your own data, and the safety side in safety-first fine-tuning explained. This guide assumes you have read the short version below and want the method.

SymptomFirst fixFine-tune when
Wrong or outdated factsRetrieval from a maintained sourceAlmost never: it is a knowledge problem
Inconsistent output formatClear instructions, a schema and worked examples in the promptThe format still drifts on a held-out set after a good prompt
Wrong label or classification on edge casesBetter examples in the prompt, then a rule-based checkYou have hundreds of reviewed edge cases and the task is stable
Tone or house styleA style guide in the system promptA style guide plus examples is not consistent enough at your volume
Cost or latency of a large model on a narrow taskA smaller model with a good promptA small tuned model beats a small prompted one and the volume justifies owning it

Which training route?

There are two practical routes. The right one depends mostly on where your data may go and on how much control you need over the finished model.

Hosted fine-tuning APIOpen weights you train yourself
Where data goesTo the provider for the training job; read their data terms firstStays on hardware you choose
What you get backA model id you can call only through that providerWeights or an adapter you can keep, move and run anywhere
EffortLow: upload a file, start a jobHigher: environment, training script, conversion, serving
Models availableWhatever the provider lists as tunable; this changesAny open-weight model whose licence allows it
Exit planYou depend on the provider for the tuned modelYou own the artefacts; hosting is your choice
  1. Step 1Define the task, the test and the baseline

    You end up with: A one-sentence task, a frozen held-out test set and a baseline score.

    Write the task in one sentence with a subject, a verb and a boundary: classify inbound support tickets into five categories; extract invoice totals into a fixed JSON schema; draft a claims summary in the house format. Then write what a correct output looks like, in enough detail that two reviewers would agree.

    Before any training, set aside a test set and do not touch it. Pick cases that include the hard and rare ones, not only the typical. A few dozen is enough to start; more makes the result steadier. Score the current prompt on it and record the number and the date. That number is the bar.

    Choose the metric to match the output: exact match for labels and fixed formats, field-level accuracy for extraction, a rubric or human review for free text. If you plan to use a model as a judge, read the pitfalls in how to validate your AI first. Decide in advance how much improvement counts as a win and what regressions disqualify the model.

  2. Step 2Build and clean the training dataset

    You end up with: A cleaned JSONL training file and a separate validation file, with provenance recorded.

    Training data is pairs of an input and the output you want. The common format is chat style: a list of messages ending with the assistant reply. Hugging Face TRL accepts conversational data like this, and also prompt and completion pairs, and applies the model chat template for you. OpenAI's supervised fine-tuning takes JSONL with a messages list in the same shape. One file can serve both routes.

    Volume: OpenAI documents a minimum of 10 examples for its service and reports improvements from 50 to 100 examples, and advises starting with about 50 good ones. For open weights, plan on a few hundred to a few thousand reviewed examples as a working range; that is our guidance, not a documented threshold, and the held-out score is the real judge. More low-quality data is worse than less clean data.

    Quality rules that matter more than volume. Cover the hard cases. Make labels consistent by settling disagreements in a written guideline. Remove duplicates and near-duplicates, which cause the model to over-weight one pattern. Remove personal data and anything you have no right to use. Record where each example came from, who approved it and when. Our page on data preparation for custom models sets out selection, sanitisation and consent.

    Keep three sets apart from the start: train, validation (used to watch for overfitting during training) and the test set from step 1. If one customer, document or conversation yields many examples, split by that unit, or near-copies leak across the boundary.

    One training example per line (chat format) · json
    {"messages": [{"role": "system", "content": "Classify the ticket as billing, bug or other. Reply with one word."}, {"role": "user", "content": "I was charged twice for March."}, {"role": "assistant", "content": "billing"}]}
    Prompt and completion pairs also work with TRL · json
    {"prompt": [{"role": "user", "content": "I was charged twice for March."}], "completion": [{"role": "assistant", "content": "billing"}]}

    Checked against: TRL SFTTrainer documentation, OpenAI supervised fine-tuning guide

  3. Step 3Route A: train through a hosted fine-tuning API

    You end up with: A fine-tuned model id you can call through the provider's API.

    Read the provider's data terms before uploading. OpenAI states that data sent to its API is not used to train or improve its models unless you opt in, and that abuse-monitoring logs are retained for up to 30 days by default. Those are statements about general API data; confirm how your training files are handled for the specific service and agreement you use, and repeat this check for any other provider.

    OpenAI's fine-tuning guide lists several methods. Supervised fine-tuning, where you provide examples of correct responses, is what this guide uses; the others (vision fine-tuning, preference optimisation and reinforcement fine-tuning) exist for different needs. At the time of reading, the page named gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14 and gpt-4.1-nano-2025-04-14 as supported for supervised fine-tuning. The tunable list changes, so check the current page before you plan around a model.

    The flow is: upload the JSONL file with purpose fine-tune, create a job with the file id and the base model, wait for it to finish, then call the resulting model id like any other model. OpenAI documents the fine-tuned model id as a string beginning ft:. Evaluate it against your test set exactly as you would any other model, and keep its id and the training file id in your model record.

    Upload the training file and start a job (OpenAI Python SDK) · python
    from openai import OpenAI
    
    client = OpenAI()  # reads OPENAI_API_KEY from the environment
    
    training_file = client.files.create(
        file=open("data/train.jsonl", "rb"),
        purpose="fine-tune",
    )
    
    job = client.fine_tuning.jobs.create(
        training_file=training_file.id,
        model="gpt-4.1-nano-2025-04-14",
    )
    print(job.id)

    Checked against: OpenAI model optimization guide, OpenAI supervised fine-tuning guide, OpenAI data controls

  4. Step 4Route B: train open weights with Hugging Face TRL

    You end up with: A LoRA adapter trained on your data that you own, ready to merge and serve.

    TRL's SFTTrainer is the standard open-source way to run supervised fine-tuning. Pass it a model id and your dataset; add a peft_config to train a small LoRA adapter instead of every weight. The TRL docs suggest a higher learning rate, around 1e-4, when training adapters, and show that conversational datasets get the chat template applied automatically.

    Choose a base model whose licence allows your use, because the tuned model inherits it. Read the licence on the exact checkpoint. Start small, in the 0.5B to 4B range, so a run takes minutes to hours. For memory sizing, QLoRA setup, merging the adapter and converting to GGUF, follow how to create your own local model: that guide has the full commands and this one stays at the method level.

    Watch validation loss. If it rises while training loss keeps falling, you are overfitting: use fewer epochs, more varied data or a lower rank. Training loss alone proves little. The only result that counts is the held-out comparison in the next step.

    Install · bash
    pip install trl peft datasets
    Minimal adapter training (train.py) · python
    from datasets import load_dataset
    from peft import LoraConfig
    from trl import SFTConfig, SFTTrainer
    
    data = load_dataset("json", data_files={"train": "data/train.jsonl", "valid": "data/valid.jsonl"})
    
    trainer = SFTTrainer(
        model="Qwen/Qwen3-4B",
        train_dataset=data["train"],
        eval_dataset=data["valid"],
        peft_config=LoraConfig(r=16, lora_alpha=16, target_modules="all-linear"),
        args=SFTConfig(
            output_dir="runs/sft",
            learning_rate=1e-4,
            num_train_epochs=2,
            eval_strategy="steps",
            eval_steps=50,
        ),
    )
    trainer.train()
    trainer.save_model("adapter-out")

    Checked against: TRL SFTTrainer documentation, PEFT LoRA developer guide

  5. Step 5Evaluate against the baseline, and against harm

    You end up with: A written verdict: ship, iterate or stop, backed by held-out numbers.

    Run the tuned model on the frozen test set from step 1, with the same prompts and settings as the baseline, and compare. Read the misses one by one: they tell you whether the data, the task definition or the base model is the limit. Compare on the slices that matter, such as the rare categories and the longest inputs, not only the average.

    Then test what you did not train on. Fine-tuning can erode safety behaviour even when the new data is harmless, so run the same refusal, over-refusal and prompt-injection checks on both models and compare. Test every language you serve. Check general ability on a few tasks unrelated to your own. Make these checks gates: a candidate that improves the task score and regresses on safety does not ship. The method for building the suite, calibrating judges and writing the evidence pack is in how to validate your AI, and evaluation and safety for custom models covers the same ground in a governed deployment.

    Record the result with the date, the model and dataset versions, the settings and who signed it off. If the improvement is smaller than the margin you set, do not ship it and call it a win. Improve the data, or go back to prompting and retrieval.

  6. Step 6Work out what it really costs, then plan ownership

    You end up with: A cost estimate that includes people and serving, and a named owner.

    Compute is the visible cost and usually the smaller one. Estimate with a formula, using your own numbers. Total cost is roughly the labelling and review hours times their rate, plus training cost, plus the cost of running the model for as long as it stays in service, plus re-testing. Training cost on a hosted API is the number of training tokens times the provider's price per million training tokens, so read the price from their pricing page on the day. On rented GPUs it is GPU hours times the hourly rate.

    A worked example with made-up numbers, to show the shape, not to quote: 1,000 examples of 600 tokens each is 600,000 tokens per pass; three passes is 1.8 million training tokens. If a provider charged a hypothetical 10 currency units per million training tokens, training would cost 18 units. If subject experts spent 60 hours reviewing at a hypothetical 50 units an hour, review costs 3,000 units. The review hours dominate the compute, which is the usual pattern.

    Then the cost that people forget: running it. A model you host needs serving capacity and monitoring, and a model on a hosted API may be priced differently from the base model. If call volume on the task is low, the people cost of owning a custom model can exceed any saving. Name an owner who will re-test it when the base model, the task or the data changes, and who can retire it.

    Cost lines to estimate (use your own figures)
    LineHow to estimate
    Labelling and reviewHours x rate, for writing, reviewing and resolving disagreements
    Evaluation harnessEngineering hours to build and maintain the test and safety suites
    Training, hosted APITraining tokens x price per million training tokens, per run, including reruns
    Training, open weightsGPU hours x hourly rate, per run, including reruns
    ServingPer-token price of the tuned model, or the cost of the server and its upkeep
    OwnershipRe-testing, monitoring, retraining and retirement over the model's life
  7. Step 7Deploy behind a stable interface and monitor it

    You end up with: The tuned model serving real traffic with a rollback path and a drift check.

    Put the model behind an interface your application already uses so you can swap it. Many servers expose an OpenAI-compatible API, so the model name and base URL can live in configuration, which is what makes a rollback to the previous model a one-line change. For serving open weights, how to self-host an LLM covers a production server, and a gateway (how to set up an LLM gateway) gives you routing and fallback.

    Roll out gradually: send a small share of traffic to the tuned model, compare against the old path on live cases, and keep the old path available. Keep the model record current: base model and revision, licence, dataset and test versions, settings, results and owner.

    Set re-test triggers: a new base model version, a change to the task or policy, a drop in the live score, or a fixed review date. Fine-tuned models are not set and forget, because the facts and rules around them keep moving.

When to stop

Stop before step 4 if the baseline already meets your threshold. Stop after step 6 if the tuned model does not beat the baseline by the margin you wrote down, or if it regresses on safety, refusals or any capability you rely on. Go back to retrieval if the misses are wrong facts. A failed run is a result: it tells you the task or the data is the problem, which is cheaper to learn before you build a deployment around the model.

Troubleshooting

What you seeLikely causeFix
The tuned model is no better than the prompted base modelThe data is too small or inconsistent, the task is really a knowledge problem, or the prompt baseline was weak.Read the misses. Fix label disagreements, add hard cases, or move to retrieval. Re-run the baseline with a stronger prompt to make sure you are comparing fairly.
Excellent on the validation file, poor on real trafficThe validation set resembles the training set too closely, or leakage between them.Rebuild the split by customer or document. Add a test set sampled from recent live inputs.
Validation loss rises while training loss fallsOverfitting.Use fewer epochs, more varied examples or a smaller adapter rank. Stop at the best validation point.
The model confidently states outdated factsFacts were trained into the weights and then changed.Move changing knowledge to retrieval. Keep fine-tuning for behaviour. See how to build a RAG system.
The model got worse at refusing unsafe requests, or over-refusesFine-tuning shifts safety behaviour even on harmless data.Run refusal and over-refusal suites on base and tuned models side by side, gate the release, and mix in general-purpose examples or reduce training strength.
The hosted job fails at validation of the training fileA malformed JSONL line, an empty message or a wrong role.Check each line parses as JSON and ends with an assistant message. Read the provider's error for the line number.
Output never terminates or repeats after local conversionChat template or end-of-turn token mismatch between training and serving.Use the same template when you serve as when you trained, and test the unconverted model first. See the local model guide.

Verify it worked

Next steps

Related guides

Frequently asked questions

How many examples do I need to fine-tune an LLM?

It depends on the task. OpenAI documents a minimum of 10 examples for its hosted service and reports gains from 50 to 100. For open weights, a few hundred to a few thousand reviewed examples is a reasonable working range. Judge by the held-out score, and favour clean, varied data over volume.

Should I use RAG or fine-tuning?

Use RAG when the model lacks knowledge that exists in your documents and changes. Use fine-tuning when it has the information but behaves inconsistently on a narrow, stable task. Many systems use a good prompt, retrieval for facts and a tuned model only for the behaviour retrieval cannot fix.

What format does fine-tuning data need to be in?

Usually JSONL with one example per line. The common shape is a messages list ending with the assistant reply, which OpenAI and Hugging Face TRL both accept. TRL also takes prompt and completion pairs, and applies the model chat template to conversational data.

Is fine-tuning safe for my data?

It depends on the route. On open weights the data stays on hardware you choose. On a hosted API it goes to the provider, so read their data terms; OpenAI states API data is not used for training unless you opt in. Remove personal data you do not need before any route.

Can fine-tuning make a model less safe?

Yes. Further training can shift refusal and safety behaviour even when the new data is harmless. Run the same safety and refusal suite on the base and the tuned model, compare them side by side and treat a regression as a failed run.

How much does fine-tuning cost?

Estimate labelling hours, evaluation work, training and serving, plus ongoing re-testing. People time usually exceeds compute. Hosted training is training tokens times the provider price per million tokens; rented GPUs are hours times the hourly rate. Step 7 has a worked example with made-up numbers.

Does Swfte offer managed fine-tuning?

No, not yet. Swfte has a Model Vault designed to hold and deploy weights you bring, but no training service. The training routes in this guide use a hosted provider API or open-source tools you run yourself.

How Swfte can help

You can follow this guide with no Swfte account. Swfte can help around the edges: keeping and deploying weights you produce, and routing between a tuned model and others behind one API.

  • Custom models: how Swfte approaches models an organisation owns
  • Domain fine-tuning: the methods compared in a governed deployment
  • Data preparation: selection, sanitisation and consent for training data
  • Connect: an OpenAI-compatible gateway to route to a tuned model and fall back to another

Swfte does not offer managed fine-tuning today, and there is no Swfte training service to buy. Where this site discusses hosting your own models, availability and plans are <custom model availability - founder to fill>. If another page on this site describes a fine-tuning flow inside Swfte, treat this guide as the current statement.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. TRL SFTTrainer documentation: dataset formats (conversational and prompt-completion), SFTTrainer, peft_config, adapter learning rate near 1e-4
  2. PEFT LoRA developer guide: LoraConfig fields r, lora_alpha and target_modules
  3. OpenAI model optimization guide: fine-tuning methods and the models listed for each, as read on 2026-10-06
  4. OpenAI supervised fine-tuning guide: JSONL messages format, 10-example minimum, gains from 50 to 100 examples, files.create and fine_tuning.jobs.create calls, ft: model id form
  5. OpenAI data controls: API data not used for training unless opted in; abuse-monitoring logs up to 30 days
  6. Swfte: when to fine-tune a model on your own data: decision ladder and readiness signals (our own post)

Topics

  • fine-tuning
  • sft
  • dataset
  • evaluation
  • openai fine-tuning
  • trl

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-fine-tune-an-llm-on-your-own-data.

Centralise your knowledge in Cortex

The desktop AI workspace: 20+ providers, local models, knowledge bases with RAG that cite their sources, MCP tools and agents, with sensitive work staying on the laptop by default.