Fine-tuning method
LoRA LLM fine-tuning: what it trains, when it fits and when it does not
Understand LoRA and QLoRA from the papers and the Hugging Face PEFT docs: what is trained, merge versus adapter serving, data needs, and when to use retrieval instead.
LoRA fine-tuning freezes the base model’s weights and trains two small low-rank matrices beside them, so you train and store far fewer parameters than full fine-tuning. QLoRA does the same on top of a 4-bit quantised base model to cut memory further. LoRA changes how a model behaves, not what it knows, so use it for format, tone and consistency, and use retrieval for facts. Swfte does not offer managed fine-tuning.
Last verified 2026-10-07. Sources are listed at the end of the page.
What does LoRA actually train?
The LoRA paper (Hu et al., 2021) describes freezing the pre-trained weights and injecting trainable rank decomposition matrices into each layer of the Transformer. The Hugging Face PEFT docs put it the same way: the weight update is represented by two smaller matrices, the original weight matrix stays frozen, and at the end both are combined.
The size of those matrices is set mainly by the rank, written r, and the shape of the original weight matrix. A higher rank gives more trainable parameters and more learning capacity. PEFT’s config also takes lora_alpha, which scales the update, and target_modules, which chooses which layers get adapters. Setting target_modules to all-linear applies it to every linear layer, as QLoRA does.
The paper’s own comparison is for GPT-3 175B against full fine-tuning with Adam: LoRA can reduce trainable parameters by 10,000 times and GPU memory by 3 times, with model quality on par or better on the models tested. That figure is the paper’s result for that setup. Your saving depends on the model, the rank and the layers you adapt.
How do LoRA, QLoRA and full fine-tuning differ?
| Method | What is trained | Memory and storage | Source |
|---|---|---|---|
| Full fine-tuning | Every weight in the model. | The most memory, and a full copy of the model per tuned version. | LoRA paper, as the baseline it compares against. |
| LoRA | Small low-rank matrices beside frozen weights. | Far fewer trainable parameters, and a small adapter file you can keep per task. | LoRA paper and PEFT docs. |
| QLoRA | LoRA adapters trained while gradients flow through a frozen 4-bit quantised base model. | The paper reports fine-tuning a 65B-parameter model on a single 48 GB GPU while preserving 16-bit fine-tuning task performance. | QLoRA paper (Dettmers et al., 2023). It introduces 4-bit NormalFloat, double quantisation and paged optimisers. |
Quantisation changes behaviour. Evaluate the exact quantised model you will serve.
Should I merge the adapter or serve it separately?
Both are documented. The choice depends on how many tasks share one base model.
| Option | What the docs say | Pick it when |
|---|---|---|
| Merge into the base model | PEFT’s merge_and_unload merges the adapter weights with the base model so the result works as a standalone model. It is not an in-place operation, so use the returned model. PEFT says LoRA adds no inference latency because adapter weights can be merged. | You have one tuned model per base, and want an ordinary full set of weights that any runtime can load. |
| Serve the adapter beside the base | vLLM serves LoRA adapters when started with --enable-lora and --lora-modules. Requests choose an adapter through the model parameter, and --max-lora-rank should be set to the highest rank you use. | Several tasks share a base model and you want one deployment to carry many small adapters. |
PEFT warns that if several adapters are active and only some are merged, the unmerged ones are silently not applied. Check which adapters are live after any merge.
When is LoRA the wrong tool?
- You need the model to know facts. Facts trained into weights go stale and are repeated with confidence. Use retrieval, as in the RAG pipeline and how to build a RAG system.
- A better prompt would do. Try instructions, a schema and worked examples first. Tune only if a held-out set still shows drift.
- Your data changes weekly. Retraining on every change is slow and hard to audit.
- You have no test set. Without one you cannot show that the adapter helped or find what it broke.
- You cannot hold the base model’s licence. An adapter and a merged model inherit the base model’s terms, so read the licence of the exact checkpoint first.
How much data does LoRA fine-tuning need?
The Hugging Face TRL and PEFT docs read for this page state no minimum number of examples. TRL documents the shapes of data it accepts: plain text, conversational messages, and prompt-completion pairs, in standard or conversational form. For conversational data it applies the chat template automatically.
The nearest documented guidance is from a hosted service, which is a different method: OpenAI’s supervised fine-tuning guide says the minimum is 10 examples, that improvements show from 50 to 100 examples, and that the right number varies greatly by use case. It recommends starting with 50 well-crafted demonstrations and evaluating the result. The QLoRA paper reports that fine-tuning on a small, high-quality dataset gave strong results.
Together these point one way: quality and coverage of reviewed examples matter more than volume. Settle label disagreements in a written guideline, remove duplicates and personal data, and keep train, validation and test sets apart. See data preparation and how to fine-tune an LLM for the method.
How do I train an adapter and decide whether to promote it?
1. Freeze the test first
Write the task in one sentence, a pass rule and a held-out set. Score the untouched model on it. That score is the baseline.
2. Train with TRL and PEFT
SFTTrainer accepts a peft_config such as LoraConfig. TRL’s docs note that adapters typically use a higher learning rate, around 1e-4, because only new parameters are learned. For QLoRA, pass a quantisation config with the PEFT config.
3. Evaluate on the held-out set
Compare against the baseline using the margin you wrote down first. Check the target task, refusals, safety behaviour and any capability you rely on.
4. Test the artefact you will serve
If you merge or quantise, re-run the evaluation on that exact file, because both can change behaviour.
5. Promote in stages
Move from development to staging to production only on a pass, and keep the previous version so you can roll back. See evaluation and safety.
Where Swfte fits
You do not need Swfte to train a LoRA adapter: TRL and PEFT run on your own hardware or a rented GPU. Swfte does not offer managed fine-tuning. Status words are Built, In progress and Roadmap.
| Capability | Status | What it means here |
|---|---|---|
| Managed fine-tuning or training service | Roadmap | Not offered. Swfte’s work on fine-tuning specialised models is research only. |
| Model Vault | Built | Hosts and versions weights you bring, with stage promotion, deploy to a dedicated endpoint and a per-model audit log. A merged model is an ordinary set of weights, which is what it is built to hold. The Vault has an asset type named LORA, but this page does not claim that adapters are served as adapters, because that is not verified. See own your model. |
| Evaluation | Built | Studio evaluation scores single-turn chat. Other evaluation is your own work. See evaluation and safety. |
| Connect gateway | Built | One OpenAI-compatible API in front of a custom endpoint, with routing, fallback, usage caps and an audit event stream. See deploy and serve. |
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- arXiv 2106.09685: LoRA: Low-Rank Adaptation of Large Language Models. Method description and the GPT-3 175B comparison.
- arXiv 2305.14314: QLoRA: Efficient Finetuning of Quantized LLMs. 4-bit base model, 65B on 48 GB, NF4, double quantisation and paged optimisers.
- Hugging Face PEFT: LoRA developer guide. LoraConfig fields, merge_and_unload, adapter switching and the partial-merge warning.
- Hugging Face PEFT: LoRA conceptual guide. How the update matrices work and merging without latency.
- Hugging Face TRL: SFT Trainer. Dataset formats, peft_config and the adapter learning rate.
- vLLM: LoRA adapters. Serving adapters with --enable-lora and --lora-modules.
- OpenAI: supervised fine-tuning guide. Example counts for a hosted service.
Frequently asked questions
What is LoRA in LLM fine-tuning?
LoRA, or Low-Rank Adaptation, freezes a pre-trained model’s weights and trains small low-rank matrices added to each layer. Only those matrices are updated, so the number of trainable parameters and the memory needed drop sharply compared with full fine-tuning. The result is a small adapter that you can keep alongside the base model or merge into it.
What is the difference between LoRA and QLoRA?
QLoRA trains LoRA adapters on top of a frozen 4-bit quantised base model, which reduces memory further. Its paper reports fine-tuning a 65B-parameter model on a single 48 GB GPU, and introduces 4-bit NormalFloat, double quantisation and paged optimisers. LoRA alone keeps the base model at higher precision, so it uses more memory.
Can I merge a LoRA adapter into the base model?
Yes. The PEFT docs describe merge_and_unload, which merges the adapter weights into the base model so the result works as a standalone model with no extra inference latency. It is not an in-place operation, so use the returned model. Re-run your evaluation on the merged file, because merging and quantising can change behaviour.
Is LoRA better than RAG for adding company knowledge?
No. LoRA changes behaviour such as format, tone and consistency, and facts trained into weights go stale. Retrieval keeps knowledge in a maintained source and updates when the source does. Use RAG for knowledge the model lacks, and LoRA for behaviour a good prompt cannot fix. Many systems use both.
How many examples do I need to fine-tune with LoRA?
The Hugging Face TRL and PEFT docs state no minimum. OpenAI’s hosted fine-tuning guide, a different method, says the minimum is 10 examples, improvements show from 50 to 100, and it recommends starting with 50 well-crafted ones. Treat that as a starting point and let a held-out set tell you whether more data helps.
Does Swfte fine-tune models with LoRA?
No. Swfte does not offer managed fine-tuning, and that stays on the Roadmap as research. You can train an adapter yourself with TRL and PEFT, merge it, and bring the weights to the Model Vault, which hosts and versions them. This page does not claim that the Vault serves adapters separately.