Short answer
You create a local model by taking a small open-weight base model, training a LoRA adapter on a few hundred to a few thousand of your own input and output examples, merging it into the base, quantising the result to GGUF and loading it in Ollama or llama.cpp. It changes behaviour and format well, and facts badly: if the problem is missing knowledge, use retrieval instead. Always compare against the untouched base on held-out cases.
The steps at a glance
- Choose a base model you may legally adapt
- Prepare and split your data
- Measure the untouched base model first
- Train a QLoRA adapter on an NVIDIA GPU
- Or train on an Apple silicon Mac with mlx-lm
- Merge the adapter into the base model
- Convert to GGUF and quantise
- Run it locally and compare with the base
- Version it, record it and plan its retirement
Before you start
Who this is for
- Engineers who want a small model that answers in a fixed format, tone or classification scheme without sending prompts to an API.
- Teams with a stable, narrow task and a few hundred reviewed examples of correct output.
- Anyone who wants to understand the whole path from data file to a model running on their own machine before paying for anything.
Probably not for you if
- Teams whose real problem is that the model does not know their documents. That is a retrieval problem; see how to build a RAG system.
- Anyone expecting a model that rivals a frontier API on open-ended work. A small adapter-tuned model is for a narrow job.
Prerequisites
- A machine you control: an Apple silicon Mac with 16 GB or more of unified memory, or a Linux or Windows machine with an NVIDIA GPU (8 GB of VRAM is the practical floor for a 4B model with QLoRA).
- Python 3.10 or newer, git, and about 30 GB of free disk for the base weights, merged copy, GGUF files and checkpoints.
- A few hundred examples of real inputs paired with outputs a qualified person has approved, and the right to train on them. Remove personal data first.
- Comfort with a terminal and reading a Python script. Nothing here needs a research background.
- Time
- About 1 day for a first working model: a few hours for data, then training time that depends on your hardware and example count.
- Cost
- Free software. You pay for your own hardware, electricity and the hours spent preparing and reviewing examples. Renting a GPU is optional.
- Hardware
- Run-only: about 5 GB of memory for an 8B model at Q4_K_M. Train: 8 to 12 GB of VRAM for a 4B model with QLoRA, or a 16 GB Apple silicon Mac for a small model with mlx-lm. See the sizing table.
- Skill
- Comfortable with Python and the command line; no machine learning background needed.
Estimates are ours, not measurements, and move with your hardware, data and network.
First decide: tune, or retrieve?
Fine-tuning moves a model's habits: output format, tone, which label to pick, how terse to be. It is a poor way to teach facts that change, because the facts end up frozen in the weights and the model states last quarter's figure with full confidence. Our own post on when to fine-tune sets out the same ladder: better prompt, then retrieval, then fine-tune.
Use this table before you spend a day on training. If your answer is in the left column, stop and use the right-hand fix.
| What is going wrong | Likely fix | Why |
|---|---|---|
| The model does not know our policies, products or documents. | Retrieval (build a RAG system) | The source of truth stays outside the weights, so updating a document updates the answer. |
| The answer is right but the format, tone or label set drifts. | A clearer prompt with worked examples first; tune if drift persists on a held-out set. | This is a behaviour problem, which an adapter handles well. |
| We run millions of near-identical calls and want a small, cheap model. | Tune a small model on outputs you have reviewed. | A narrow task with high volume is where a small tuned model pays back. |
| We have fewer than about a hundred reviewed examples. | Write the examples into the prompt and collect more data. | Too little data teaches the model your noise, not your task. |
What your hardware can do: honest sizing
Memory is the limit that decides everything. A model with N billion parameters stored in 16-bit precision needs about 2 bytes per parameter, so roughly 2N gigabytes for the weights alone. Four-bit quantisation brings that to a little over 0.5 bytes per parameter. The llama.cpp quantisation notes give a measured example: Llama 3.1 8B is 14.96 GiB at F16 and 4.58 GiB at Q4_K_M. Training adds optimiser state, activations and the context length on top, so the training column below is a planning estimate, not a measurement.
| Task | Model size | Memory to plan for | Basis |
|---|---|---|---|
| Run only, Q4_K_M | 8B | About 5 GB plus context | llama.cpp notes: 4.58 GiB for Llama 3.1 8B at Q4_K_M |
| Run only, Q4_K_M | 4B | About 3 GB plus context | Arithmetic: 4B parameters x about 0.6 bytes |
| QLoRA training | 4B | 8 to 12 GB of VRAM | Estimate: 4-bit weights about 2 to 3 GB, plus adapter state, activations and context |
| LoRA training in 16-bit | 4B | 16 GB or more | Arithmetic: weights alone are about 8 GB, plus training state |
| LoRA training in 16-bit | 8B | 24 GB or more | Arithmetic: weights alone are about 16 GB, plus training state |
| Smallest QLoRA run | Under 1B | About 3 GB | Unsloth fine-tuning guide: as little as 3 GB of VRAM with QLoRA |
Step 1Choose a base model you may legally adapt
You end up with: One named base model, its licence recorded, small enough for your machine.
Pick an instruction-tuned open-weight model in the 0.5B to 4B range for a first run. Smaller models train faster, fit on modest hardware and make mistakes you can see quickly. Read the licence on the exact checkpoint you will use, because your tuned model inherits it.
The table shows what the Hugging Face model API reported for these checkpoints on 6 October 2026. Gated means you must accept terms before you can download. If a licence is anything other than Apache-2.0 or MIT, read the licence text itself before you build on it, and ask your legal team if the model will be sold or shipped to customers.
This guide uses
Qwen/Qwen3-4Bin its examples, which is listed under Apache-2.0. Swap in any model whose architecture your chosen tool supports; check the tool documentation, not the model card, for that.Hugging Face model Parameters Licence tag Gated Qwen/Qwen3-0.6B 0.75 billion apache-2.0 No HuggingFaceTB/SmolLM3-3B 3.08 billion apache-2.0 No microsoft/Phi-4-mini-instruct 3.84 billion mit No Qwen/Qwen3-4B 4.02 billion apache-2.0 No Qwen/Qwen3-8B 8.19 billion apache-2.0 No meta-llama/Llama-3.1-8B-Instruct 8.03 billion llama3.1 (custom licence) Yes, manual approval Checked against: Hugging Face model API and model cards
Step 2Prepare and split your data
You end up with: Three JSONL files: train, valid and test, with no overlap between them.
Write each example as one line of JSON in chat format: a list of messages with the final one from the assistant. This is the conversational format both TRL and mlx-lm read, so the same files work on every path in this guide.
Quality beats quantity. Include the hard cases and the inputs that currently fail, not only easy ones. If two reviewers would label the same input differently, settle the rule in writing before training, or the model learns the disagreement. Strip personal data and anything you have no right to train on.
Split before you do anything else, and never look at the test file while tuning. The script below shuffles with a fixed seed and writes an 80, 10, 10 split. It is our script, not from any tool documentation. If the same customer, document or conversation appears in several examples, split by that unit instead, otherwise near-duplicates leak from train into test and flatter the result.
One example per line (data/all.jsonl) · json {"messages": [{"role": "system", "content": "Classify the support ticket as billing, bug or other. Reply with one word."}, {"role": "user", "content": "I was charged twice for March."}, {"role": "assistant", "content": "billing"}]}Split into train, valid and test (split.py) · python import json, random, pathlib random.seed(0) rows = [json.loads(line) for line in open("data/all.jsonl", encoding="utf-8") if line.strip()] random.shuffle(rows) n = len(rows) cut_train, cut_valid = int(n * 0.8), int(n * 0.9) parts = {"train": rows[:cut_train], "valid": rows[cut_train:cut_valid], "test": rows[cut_valid:]} out = pathlib.Path("data") for name, part in parts.items(): with open(out / f"{name}.jsonl", "w", encoding="utf-8") as f: for row in part: f.write(json.dumps(row, ensure_ascii=False) + "\n") print(name, len(part))Run it · bash python split.pyExample output for 1,000 examples
train 800 valid 100 test 100Checked against: mlx-lm LORA.md, TRL SFTTrainer documentation
Step 3Measure the untouched base model first
You end up with: A baseline score on your test set that the tuned model must beat.
Without a baseline you cannot say the tuning helped. Install Ollama, pull the same base model, and score it on the test file. On macOS and Linux the install is one line; on Windows use the PowerShell line. Ollama serves an OpenAI-compatible API on port 11434, so a short script can send each test prompt and compare the reply with the reference.
The script uses exact match, which suits classification and fixed formats. For free text, replace the comparison with a rubric or a human review; the guide how to validate your AI covers that, including the traps of using a model as the judge. If your model prints reasoning text before the answer, strip it before comparing. Keep temperature at zero so reruns are comparable.
Record the number and the date. If the base model already clears your bar, stop here: a better prompt is cheaper than a tuned model you now have to own.
Install Ollama (macOS or Linux) · bash curl -fsSL https://ollama.com/install.sh | shInstall Ollama (Windows PowerShell) · powershell irm https://ollama.com/install.ps1 | iexPull the base model · bash ollama pull qwen3:4bScore a model on the test set (evaluate.py) · python import json, sys from openai import OpenAI client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama") # key is required but ignored model, path = sys.argv[1], sys.argv[2] hits = total = 0 for line in open(path, encoding="utf-8"): if not line.strip(): continue messages = json.loads(line)["messages"] prompt, gold = messages[:-1], messages[-1]["content"].strip() reply = client.chat.completions.create(model=model, messages=prompt, temperature=0) answer = reply.choices[0].message.content.strip() hits += int(answer == gold) total += 1 if answer != gold: print("MISS | gold:", gold[:60], "| got:", answer[:60]) print(f"{model}: {hits}/{total} exact matches")Score the base model · bash pip install openai python evaluate.py qwen3:4b data/test.jsonlChecked against: Ollama README and download page, Ollama OpenAI compatibility, Ollama model library: qwen3
Step 4Train a QLoRA adapter on an NVIDIA GPU
You end up with: A LoRA adapter saved in adapter-out/, trained on your train file and checked against valid.
QLoRA loads the base model in 4-bit and trains small adapter matrices on top, which is why it fits in far less memory than training in 16-bit. The Hugging Face docs describe it as quantising the model to 4 bits and then training it with LoRA. TRL, PEFT and bitsandbytes are the standard open-source stack. The script below follows the documented TRL pattern: pass a
peft_configand aquantization_configtoSFTTrainer, usetarget_modules="all-linear"as the PEFT quantisation guide advises for QLoRA, and use a learning rate near 1e-4, which the TRL docs suggest for adapters.Treat the epoch count, batch size and rank as starting points, not truths. Watch the validation loss. If it falls and then climbs while training loss keeps falling, you are overfitting: stop earlier, use fewer epochs or add data. Unsloth's guide recommends one to three epochs.
If you prefer a faster, lower-memory wrapper, Unsloth installs with
uv pip install unsloth --torch-backend=autoand publishes notebooks for specific models. Axolotl is configuration-driven: you write one YAML file and runaxolotl train. Both need an NVIDIA GPU for this path; Axolotl asks for Ampere or newer for bf16 and Flash Attention.Install the libraries (Python virtual environment) · bash pip install --upgrade transformers accelerate bitsandbytes pip install trl peft datasetsTrain (train_qlora.py) · python import torch from datasets import load_dataset from peft import LoraConfig from transformers import BitsAndBytesConfig from trl import SFTConfig, SFTTrainer BASE = "Qwen/Qwen3-4B" data = load_dataset("json", data_files={"train": "data/train.jsonl", "valid": "data/valid.jsonl"}) bnb = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) trainer = SFTTrainer( model=BASE, train_dataset=data["train"], eval_dataset=data["valid"], quantization_config=bnb, peft_config=LoraConfig(r=16, lora_alpha=16, target_modules="all-linear"), args=SFTConfig( output_dir="runs/qlora", learning_rate=1e-4, num_train_epochs=2, per_device_train_batch_size=2, gradient_accumulation_steps=8, assistant_only_loss=True, eval_strategy="steps", eval_steps=50, logging_steps=10, ), ) trainer.train() trainer.save_model("adapter-out")Run it · bash python train_qlora.pyAlternative: Axolotl (config-driven) · bash axolotl fetch examples axolotl train examples/llama-3/lora-1b.ymlChecked against: TRL SFTTrainer documentation, Transformers bitsandbytes guide, PEFT quantization guide, Unsloth fine-tuning guide, Unsloth README, Axolotl README, Axolotl getting started
Step 5Or train on an Apple silicon Mac with mlx-lm
You end up with: A LoRA adapter in adapters/ and a test-set perplexity figure.
On a Mac, mlx-lm is the simplest route. It trains LoRA by default, and if you point it at a quantised model it trains QLoRA. Your three JSONL files go in one directory named by
--data, with the namestrain.jsonl,valid.jsonlandtest.jsonl.The mlx-lm notes list the model families its LoRA training supports (Mistral, Llama, Phi2, Mixtral, Qwen2, Gemma, OLMo, MiniCPM, InternLM2 at the time of reading). If your base model is newer than that list, run
mlx_lm.lora --helpand try a short run before committing a day to it; if the architecture is not supported the command will tell you early.To cut memory, lower
--batch-size(default 4), lower--num-layers(default 16), add--grad-checkpoint, or shorten examples. The same notes report about 250 tokens per second on an M1 Max with 32 GB for a 7B model with batch size 1 and 4 layers on a sample dataset. Use that only as a rough guide: one million training tokens at 250 tokens per second is about 67 minutes. Run a few steps and read your own rate.Install · bash pip install "mlx-lm[train]"Train · bash mlx_lm.lora \ --model <your-base-model> \ --train \ --data data \ --iters 600 \ --mask-prompt \ --adapter-path adaptersTest-set perplexity with the adapter · bash mlx_lm.lora \ --model <your-base-model> \ --adapter-path adapters \ --data data \ --testTry the adapter on one prompt · bash mlx_lm.generate \ --model <your-base-model> \ --adapter-path adapters \ --prompt "I was charged twice for March."Checked against: mlx-lm LORA.md, mlx-lm README
Step 6Merge the adapter into the base model
You end up with: A single full model directory (merged/) with your changes baked in.
A LoRA adapter is a small file that sits on top of the base. To convert to GGUF you first need one ordinary model directory, so merge the adapter into the base weights. With PEFT this is
merge_and_unload(). The PEFT docs are explicit that it is not an in-place operation: you must keep the model it returns.Load the base in 16-bit for the merge, not in 4-bit, so the merged weights are not rounded twice. This needs enough memory for the full-precision model, which is about 8 GB for 4B parameters, but it can run on CPU if the GPU is too small.
Other tools have their own merge step. Axolotl uses
axolotl merge-lorawith the same YAML and the adapter directory. Unsloth hassave_pretrained_mergedwithsave_method="merged_16bit". On a Mac,mlx_lm.fuse --model <your-base-model>loads adapters fromadapters/and writesfused_model/; its GGUF export is limited to Mistral, Mixtral and Llama-style models in fp16, so for other families run the fused model with mlx-lm or train on an NVIDIA machine for the GGUF route.Merge with PEFT (merge.py) · python import torch from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer BASE = "Qwen/Qwen3-4B" base = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16) model = PeftModel.from_pretrained(base, "adapter-out") model = model.merge_and_unload() # returns the merged model; not in place model.save_pretrained("merged") AutoTokenizer.from_pretrained(BASE).save_pretrained("merged")Run it · bash python merge.pyAlternative: Axolotl · bash axolotl merge-lora train_config.yml --lora-model-dir="./outputs/lora-out"Alternative: mlx-lm · bash mlx_lm.fuse --model <your-base-model>Checked against: PEFT LoRA developer guide, Axolotl getting started, Unsloth: saving to GGUF, mlx-lm LORA.md
Step 7Convert to GGUF and quantise
You end up with: A single .gguf file around a quarter the size of the 16-bit model.
GGUF is the file format llama.cpp and Ollama read. Conversion is two phases: convert the merged Hugging Face model to a high-precision GGUF, then quantise that file. Quantising from 16-bit or 32-bit gives better quality than re-quantising something already quantised.
Get llama.cpp from its repository and build it with CMake. Install its Python requirements, run
convert_hf_to_gguf.pyon the merged directory, then runllama-quantizewith the method name.Q4_K_Mis the usual first choice. The llama.cpp notes show Llama 3.1 8B at 14.96 GiB in F16 and 4.58 GiB at Q4_K_M, and 7.95 GiB at Q8_0. Their measured perplexity table is how you would pick a different method if Q4_K_M loses too much quality on your task.The llama.cpp notes warn that some newer models need a newer
transformersthan its requirements file installs, and saypip install -U transformersis safe. If conversion fails with an unknown architecture, update llama.cpp before anything else.Unsloth users can skip the manual route:
model.save_pretrained_gguf("directory", tokenizer, quantization_method="q4_k_m")does the merge and quantise in one call. Unsloth also warns that the most common cause of gibberish is a different chat template at inference than at training.Get and build llama.cpp · bash git clone https://github.com/ggml-org/llama.cpp cd llama.cpp cmake -B build cmake --build build --config Release python3 -m pip install -r requirements.txtConvert the merged model to GGUF (bf16) · bash python convert_hf_to_gguf.py ../merged --outfile my-model-bf16.gguf --outtype bf16Quantise to Q4_K_M · bash ./build/bin/llama-quantize my-model-bf16.gguf my-model-Q4_K_M.gguf Q4_K_MChecked against: llama.cpp quantize README, llama.cpp build guide, Unsloth: saving to GGUF
Step 8Run it locally and compare with the base
You end up with: The tuned model answering in Ollama, with a before and after score on the held-out test set.
Ollama does not quantise GGUF files on import, which is why you did it in the previous step. Create a
Modelfilethat points at the file, build the model withollama create, and then run the sameevaluate.pyagainst it. Compare the number with the baseline from step 3. If the tuned model beats the base by the margin you set in advance, and nothing else you care about got worse, it passes.If you would rather use llama.cpp directly,
llama-server -m my-model-Q4_K_M.ggufstarts an OpenAI-compatible server; the server notes say it listens on 127.0.0.1:8080 by default. Point the evaluation script athttp://localhost:8080/v1/to score it the same way.Also test what you did not train on. Ask general questions, ask things you want it to refuse, and try inputs in other languages if you serve them. Our post on safety-first fine-tuning lists what to check. For a fuller method covering judges, regression gates and an evidence pack, see how to validate your AI.
Modelfile · dockerfile FROM ./my-model-Q4_K_M.ggufBuild and try the model · bash ollama create my-model ollama run my-model "I was charged twice for March."Score the tuned model · bash python evaluate.py my-model data/test.jsonlAlternative: llama.cpp server · bash ./build/bin/llama-server -m my-model-Q4_K_M.ggufWhat a passing comparison looks like (numbers are illustrative)
qwen3:4b: 61/100 exact matches my-model: 91/100 exact matchesChecked against: Ollama: importing a model, Ollama CLI reference, llama.cpp server README
Step 9Version it, record it and plan its retirement
You end up with: A short model record you can hand to a reviewer, and a rule for when to retrain.
A tuned model is a new artefact someone has to own. Write down: the base model and exact revision, its licence and the date you read it, the dataset version and how examples were approved, training settings, the quantisation method, the baseline and tuned scores, and the safety checks you ran. Store the adapter, the merged weights and the GGUF together with a hash of each.
Decide when to retrain: when the task definition changes, when the base model is superseded, or when monitoring shows the score slipping on fresh cases. Decide when to retire it, and who is allowed to say so.
If you need to serve it to more than one person, put it behind a proper server rather than a laptop. How to self-host an LLM covers that. Swfte does not offer managed fine-tuning today; its Model Vault is designed to hold weights you bring, with versions and promotion stages. Details are on the custom models pages linked below.
When to stop and use RAG instead
Stop after step 4 if the baseline already passes your threshold: a better prompt may be all you need. Stop after step 6 if the tuned model does not beat the base on the held-out set by a margin you decided in advance. Move to retrieval if the failures on the test set are wrong facts rather than wrong format. Move to a larger base model only after you have a clean evaluation, because otherwise you cannot tell whether the bigger model helped.
If the tuned model improves the task score but degrades something else you care about, such as refusing unsafe requests or answering general questions, treat that as a failed run. Fine-tuning can erode safety behaviour even on harmless data; our post on safety-first fine-tuning explains why and what to test.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| CUDA out of memory during training | Batch size, sequence length or model size is too large for your VRAM. | Lower per_device_train_batch_size and raise gradient_accumulation_steps by the same factor, shorten examples, or use a smaller base model. Confirm 4-bit loading is on. |
| Training loss falls but validation loss rises | Overfitting to a small dataset. | Stop earlier, reduce epochs, add more varied examples, or lower the learning rate. Do not judge on training loss. |
| The tuned model repeats itself or never stops, or answers in gibberish after conversion | The chat template at inference differs from the one used at training (Unsloth documents this as the most common cause), or the end-of-turn token was not aligned. | Use the same chat template when you run it as when you trained. Test the merged Hugging Face model first; if it is fine there, the problem is in conversion or the Modelfile template. |
| convert_hf_to_gguf.py fails with an unsupported architecture or tokenizer error | llama.cpp or transformers is older than the model. | Update the llama.cpp checkout and run pip install -U transformers, which the llama.cpp notes say is safe. |
| The tuned model scores no better than the base | Too few or inconsistent examples, a task that is really a knowledge problem, or data leakage that made the baseline look worse than it is. | Read the misses. If they are wrong facts, switch to retrieval. If labels disagree, fix the guideline and relabel. Check train and test do not overlap. |
| mlx_lm.lora errors on your base model | The architecture is not supported by the installed mlx-lm version. | Run pip install -U mlx-lm, check mlx_lm.lora --help, or choose a model from the supported families listed in the mlx-lm notes. |
| The model got better at the task but worse at declining unsafe requests | Fine-tuning can erode safety behaviour even on harmless data. | Run a refusal and over-refusal suite on base and tuned side by side and gate the release on it. Try a lower rank, fewer epochs, or mix in general examples. |
Verify it worked
Next steps
- How to validate your AI: build a bigger evaluation set, gate releases and keep an evidence pack
- How to run LLMs locally: get more from Ollama, llama.cpp and LM Studio, and choose quantisation levels
- How to self-host an LLM: serve the model to a team on a GPU server instead of one laptop
- How to build a RAG system: give any model, tuned or not, access to documents that change
Related guides
- How to Validate Your AI: Eval Sets, Gates, Evidence: A system-level method to validate an AI product: define the task and risk, build a held-out eval set, score it, gate releases, sample for human review, monitor and keep an evidence pack.
- How to Run LLMs Locally: Ollama, LM Studio, llama.cpp: Install Ollama, LM Studio or llama.cpp, download a model that fits your memory, chat with it and call it from code through a local OpenAI-compatible endpoint.
- How to Fine-Tune an LLM on Your Own Data (2026): The decision and the method: when fine-tuning beats prompting and retrieval, how to build and split the dataset, two training routes, how to evaluate, and what it really costs.
- How to Self-Host an LLM with vLLM (2026 Guide): Serve an open-weight model as a private, OpenAI-compatible endpoint on your own GPU server, with memory sizing, authentication, TLS, metrics and an upgrade routine.
- How to Build a RAG System: Step-by-Step, Runs Locally: Build a retrieval-augmented generation system end to end on one machine, with hybrid search, per-group permissions, citations and a retrieval test set.
Frequently asked questions
Can I train my own LLM on a laptop?
You can adapt a small open-weight model on a laptop, not train one from scratch. With QLoRA or mlx-lm, models in roughly the 0.5B to 4B range are realistic on a 16 GB machine. Training a new foundation model needs data and compute far beyond a laptop.
How much data do I need to fine-tune a model?
It depends on the task. OpenAI documents a floor of 10 examples and sees gains from 50 to 100 for its hosted service. For a local adapter, plan on a few hundred reviewed examples to start, and judge by the held-out score, not by a number.
What is the difference between LoRA and QLoRA?
LoRA trains small adapter matrices while the base weights stay frozen. QLoRA does the same with the base model loaded in 4-bit, which cuts memory a lot. Unsloth describes it as saving about 75 per cent of memory against 16-bit.
Is fine-tuning better than RAG?
They solve different problems. Fine-tuning changes behaviour: format, tone, narrow classification. RAG supplies knowledge that changes. If answers are wrong because the model lacks facts, use retrieval first.
What is a GGUF file?
GGUF is the model file format read by llama.cpp and Ollama. It holds the weights, usually quantised to 4 to 8 bits, plus tokeniser metadata in one file, so a model can be copied and loaded on a laptop or server.
Which quantisation should I use?
Start with Q4_K_M. In the llama.cpp measurements for Llama 3.1 8B it is 4.58 GiB against 14.96 GiB for F16. If it hurts your task score, try Q5_K_M or Q8_0 and keep the bf16 GGUF so you can requantise.
Does Swfte offer managed fine-tuning?
No, not today. Swfte has a Model Vault designed to hold and deploy weights you bring. The training steps in this guide run on your own hardware or rented GPUs, using open-source tools.
How Swfte can help
You can do every step above without Swfte. If you want somewhere to keep, version and deploy the weights you produce, or a desktop app to chat with local models, these pages describe what exists.
- Custom models: how Swfte approaches models you own
- Deploy and serve: serving weights you bring
- Deploy models: self-hosting guides for open-weight models
- Open-source model testing: how we test open-weight models before recommending them
- Cortex: a desktop app that runs local models through Ollama or LM Studio
Swfte does not offer managed fine-tuning yet: there is no Swfte training service to run for you. Availability and plans for hosting custom models: <custom model availability - founder to fill>.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- Hugging Face model API and model cards: licence tag, parameter count and gating of the models in the table, read on 2026-10-06
- TRL SFTTrainer documentation: dataset formats, SFTTrainer with peft_config and quantization_config, assistant_only_loss, adapter learning rate
- PEFT LoRA developer guide: LoraConfig fields and merge_and_unload() not being in place
- PEFT quantization guide: QLoRA definition and target_modules all-linear
- Transformers bitsandbytes guide: BitsAndBytesConfig 4-bit and nf4 settings, install line
- mlx-lm LORA.md: mlx_lm.lora, data layout, test, generate, fuse, memory tips, GGUF export limits
- mlx-lm README: pip install mlx-lm
- Unsloth README: uv pip install unsloth --torch-backend=auto, supported platforms, licence
- Unsloth fine-tuning guide: QLoRA memory saving, minimum VRAM, epochs, learning rate
- Unsloth: saving to GGUF: save_pretrained_gguf, save_pretrained_merged, chat template warning
- Axolotl README: requirements (NVIDIA GPU, Ampere or newer for bf16), axolotl fetch examples and train
- Axolotl getting started: adapter: qlora, axolotl merge-lora, inference commands
- llama.cpp quantize README: convert_hf_to_gguf.py, llama-quantize Q4_K_M, size table for Llama 3.1 8B
- llama.cpp build guide: cmake -B build and cmake --build build --config Release
- llama.cpp README: install options, and the newer llama cli and llama serve wrapper commands
- llama.cpp server README: llama-server default address 127.0.0.1:8080, OpenAI-compatible chat completions
- Ollama README and download page: install commands for macOS, Linux and Windows
- Ollama CLI reference: ollama run, pull, create, ps
- Ollama: importing a model: Modelfile FROM a GGUF, ollama create, no quantisation on import
- Ollama OpenAI compatibility: base_url http://localhost:11434/v1/, api key ignored
- OpenAI supervised fine-tuning guide: minimum of 10 examples and gains from 50 to 100 for the hosted service
- Ollama model library: qwen3: qwen3:4b tag exists
Topics
- lora
- qlora
- gguf
- ollama
- llama.cpp
- mlx
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-create-your-own-local-model.