# How to create your own local model

Canonical: https://www.swfte.com/how-to-create-your-own-local-model
Last verified: 2026-10-06
Difficulty: Advanced
Time: About 1 day for a first working model: a few hours for data, then training time that depends on your hardware and example count.
Cost: Free software. You pay for your own hardware, electricity and the hours spent preparing and reviewing examples. Renting a GPU is optional.
Hardware: Run-only: about 5 GB of memory for an 8B model at Q4_K_M. Train: 8 to 12 GB of VRAM for a 4B model with QLoRA, or a 16 GB Apple silicon Mac for a small model with mlx-lm. See the sizing table.

## Short answer

You create a local model by taking a small open-weight base model, training a LoRA adapter on a few hundred to a few thousand of your own input and output examples, merging it into the base, quantising the result to GGUF and loading it in Ollama or llama.cpp. It changes behaviour and format well, and facts badly: if the problem is missing knowledge, use retrieval instead. Always compare against the untouched base on held-out cases.

## Who this is for

- Engineers who want a small model that answers in a fixed format, tone or classification scheme without sending prompts to an API.
- Teams with a stable, narrow task and a few hundred reviewed examples of correct output.
- Anyone who wants to understand the whole path from data file to a model running on their own machine before paying for anything.

Not for:
- Teams whose real problem is that the model does not know their documents. That is a retrieval problem; see [how to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system).
- Anyone expecting a model that rivals a frontier API on open-ended work. A small adapter-tuned model is for a narrow job.

## Prerequisites

- A machine you control: an Apple silicon Mac with 16 GB or more of unified memory, or a Linux or Windows machine with an NVIDIA GPU (8 GB of VRAM is the practical floor for a 4B model with QLoRA).
- Python 3.10 or newer, git, and about 30 GB of free disk for the base weights, merged copy, GGUF files and checkpoints.
- A few hundred examples of real inputs paired with outputs a qualified person has approved, and the right to train on them. Remove personal data first.
- Comfort with a terminal and reading a Python script. Nothing here needs a research background.

## First decide: tune, or retrieve?

Fine-tuning moves a model's habits: output format, tone, which label to pick, how terse to be. It is a poor way to teach facts that change, because the facts end up frozen in the weights and the model states last quarter's figure with full confidence. Our own post on [when to fine-tune](https://www.swfte.com/blog/when-to-fine-tune-a-model-on-your-own-data-2026) sets out the same ladder: better prompt, then retrieval, then fine-tune.

Use this table before you spend a day on training. If your answer is in the left column, stop and use the right-hand fix.

| What is going wrong | Likely fix | Why |
| --- | --- | --- |
| The model does not know our policies, products or documents. | Retrieval ([build a RAG system](https://www.swfte.com/how-to-build-a-rag-system)) | The source of truth stays outside the weights, so updating a document updates the answer. |
| The answer is right but the format, tone or label set drifts. | A clearer prompt with worked examples first; tune if drift persists on a held-out set. | This is a behaviour problem, which an adapter handles well. |
| We run millions of near-identical calls and want a small, cheap model. | Tune a small model on outputs you have reviewed. | A narrow task with high volume is where a small tuned model pays back. |
| We have fewer than about a hundred reviewed examples. | Write the examples into the prompt and collect more data. | Too little data teaches the model your noise, not your task. |

## What your hardware can do: honest sizing

Memory is the limit that decides everything. A model with N billion parameters stored in 16-bit precision needs about 2 bytes per parameter, so roughly 2N gigabytes for the weights alone. Four-bit quantisation brings that to a little over 0.5 bytes per parameter. The llama.cpp quantisation notes give a measured example: Llama 3.1 8B is 14.96 GiB at F16 and 4.58 GiB at Q4_K_M. Training adds optimiser state, activations and the context length on top, so the training column below is a planning estimate, not a measurement.

**Planning estimates. Weights-only figures come from simple arithmetic or the llama.cpp notes; training headroom is our estimate.**

| Task | Model size | Memory to plan for | Basis |
| --- | --- | --- | --- |
| Run only, Q4_K_M | 8B | About 5 GB plus context | llama.cpp notes: 4.58 GiB for Llama 3.1 8B at Q4_K_M |
| Run only, Q4_K_M | 4B | About 3 GB plus context | Arithmetic: 4B parameters x about 0.6 bytes |
| QLoRA training | 4B | 8 to 12 GB of VRAM | Estimate: 4-bit weights about 2 to 3 GB, plus adapter state, activations and context |
| LoRA training in 16-bit | 4B | 16 GB or more | Arithmetic: weights alone are about 8 GB, plus training state |
| LoRA training in 16-bit | 8B | 24 GB or more | Arithmetic: weights alone are about 16 GB, plus training state |
| Smallest QLoRA run | Under 1B | About 3 GB | Unsloth fine-tuning guide: as little as 3 GB of VRAM with QLoRA |

> TIP: Start with the smallest model that could plausibly do the job. A 0.6B to 4B model that trains in an hour teaches you more about your data than an 8B model that runs out of memory at step 40.

## Steps

### Step 1: Choose a base model you may legally adapt

Outcome: One named base model, its licence recorded, small enough for your machine.

Pick an instruction-tuned open-weight model in the 0.5B to 4B range for a first run. Smaller models train faster, fit on modest hardware and make mistakes you can see quickly. Read the licence on the exact checkpoint you will use, because your tuned model inherits it.

The table shows what the Hugging Face model API reported for these checkpoints on 6 October 2026. Gated means you must accept terms before you can download. If a licence is anything other than Apache-2.0 or MIT, read the licence text itself before you build on it, and ask your legal team if the model will be sold or shipped to customers.

This guide uses `Qwen/Qwen3-4B` in its examples, which is listed under Apache-2.0. Swap in any model whose architecture your chosen tool supports; check the tool documentation, not the model card, for that.

| Hugging Face model | Parameters | Licence tag | Gated |
| --- | --- | --- | --- |
| Qwen/Qwen3-0.6B | 0.75 billion | apache-2.0 | No |
| HuggingFaceTB/SmolLM3-3B | 3.08 billion | apache-2.0 | No |
| microsoft/Phi-4-mini-instruct | 3.84 billion | mit | No |
| Qwen/Qwen3-4B | 4.02 billion | apache-2.0 | No |
| Qwen/Qwen3-8B | 8.19 billion | apache-2.0 | No |
| meta-llama/Llama-3.1-8B-Instruct | 8.03 billion | llama3.1 (custom licence) | Yes, manual approval |

> NOTE: Licence tags are a quick filter, not legal advice. Open the licence file shipped with the weights and save a copy with the date you read it.

### Step 2: Prepare and split your data

Outcome: Three JSONL files: train, valid and test, with no overlap between them.

Write each example as one line of JSON in chat format: a list of messages with the final one from the assistant. This is the conversational format both TRL and mlx-lm read, so the same files work on every path in this guide.

Quality beats quantity. Include the hard cases and the inputs that currently fail, not only easy ones. If two reviewers would label the same input differently, settle the rule in writing before training, or the model learns the disagreement. Strip personal data and anything you have no right to train on.

Split before you do anything else, and never look at the test file while tuning. The script below shuffles with a fixed seed and writes an 80, 10, 10 split. It is our script, not from any tool documentation. If the same customer, document or conversation appears in several examples, split by that unit instead, otherwise near-duplicates leak from train into test and flatter the result.

One example per line (data/all.jsonl):

```json
{"messages": [{"role": "system", "content": "Classify the support ticket as billing, bug or other. Reply with one word."}, {"role": "user", "content": "I was charged twice for March."}, {"role": "assistant", "content": "billing"}]}
```

Split into train, valid and test (split.py):

```python
import json, random, pathlib

random.seed(0)
rows = [json.loads(line) for line in open("data/all.jsonl", encoding="utf-8") if line.strip()]
random.shuffle(rows)

n = len(rows)
cut_train, cut_valid = int(n * 0.8), int(n * 0.9)
parts = {"train": rows[:cut_train], "valid": rows[cut_train:cut_valid], "test": rows[cut_valid:]}

out = pathlib.Path("data")
for name, part in parts.items():
    with open(out / f"{name}.jsonl", "w", encoding="utf-8") as f:
        for row in part:
            f.write(json.dumps(row, ensure_ascii=False) + "\n")
    print(name, len(part))
```

Run it:

```bash
python split.py
```

Example output for 1,000 examples:

```text
train 800
valid 100
test 100
```

### Step 3: Measure the untouched base model first

Outcome: A baseline score on your test set that the tuned model must beat.

Without a baseline you cannot say the tuning helped. Install Ollama, pull the same base model, and score it on the test file. On macOS and Linux the install is one line; on Windows use the PowerShell line. Ollama serves an OpenAI-compatible API on port 11434, so a short script can send each test prompt and compare the reply with the reference.

The script uses exact match, which suits classification and fixed formats. For free text, replace the comparison with a rubric or a human review; the guide [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai) covers that, including the traps of using a model as the judge. If your model prints reasoning text before the answer, strip it before comparing. Keep temperature at zero so reruns are comparable.

Record the number and the date. If the base model already clears your bar, stop here: a better prompt is cheaper than a tuned model you now have to own.

Install Ollama (macOS or Linux):

```bash
curl -fsSL https://ollama.com/install.sh | sh
```

Install Ollama (Windows PowerShell):

```powershell
irm https://ollama.com/install.ps1 | iex
```

Pull the base model:

```bash
ollama pull qwen3:4b
```

Score a model on the test set (evaluate.py):

```python
import json, sys
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")  # key is required but ignored
model, path = sys.argv[1], sys.argv[2]

hits = total = 0
for line in open(path, encoding="utf-8"):
    if not line.strip():
        continue
    messages = json.loads(line)["messages"]
    prompt, gold = messages[:-1], messages[-1]["content"].strip()
    reply = client.chat.completions.create(model=model, messages=prompt, temperature=0)
    answer = reply.choices[0].message.content.strip()
    hits += int(answer == gold)
    total += 1
    if answer != gold:
        print("MISS | gold:", gold[:60], "| got:", answer[:60])

print(f"{model}: {hits}/{total} exact matches")
```

Score the base model:

```bash
pip install openai
python evaluate.py qwen3:4b data/test.jsonl
```

> TIP: Read the MISS lines. If most misses are wrong facts, retrieval will help more than tuning. If most are formatting, tuning is the right tool.

### Step 4: Train a QLoRA adapter on an NVIDIA GPU

Outcome: A LoRA adapter saved in adapter-out/, trained on your train file and checked against valid.

QLoRA loads the base model in 4-bit and trains small adapter matrices on top, which is why it fits in far less memory than training in 16-bit. The Hugging Face docs describe it as quantising the model to 4 bits and then training it with LoRA. TRL, PEFT and bitsandbytes are the standard open-source stack. The script below follows the documented TRL pattern: pass a `peft_config` and a `quantization_config` to `SFTTrainer`, use `target_modules="all-linear"` as the PEFT quantisation guide advises for QLoRA, and use a learning rate near 1e-4, which the TRL docs suggest for adapters.

Treat the epoch count, batch size and rank as starting points, not truths. Watch the validation loss. If it falls and then climbs while training loss keeps falling, you are overfitting: stop earlier, use fewer epochs or add data. Unsloth's guide recommends one to three epochs.

If you prefer a faster, lower-memory wrapper, Unsloth installs with `uv pip install unsloth --torch-backend=auto` and publishes notebooks for specific models. Axolotl is configuration-driven: you write one YAML file and run `axolotl train`. Both need an NVIDIA GPU for this path; Axolotl asks for Ampere or newer for bf16 and Flash Attention.

Install the libraries (Python virtual environment):

```bash
pip install --upgrade transformers accelerate bitsandbytes
pip install trl peft datasets
```

Train (train_qlora.py):

```python
import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

BASE = "Qwen/Qwen3-4B"

data = load_dataset("json", data_files={"train": "data/train.jsonl", "valid": "data/valid.jsonl"})

bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)

trainer = SFTTrainer(
    model=BASE,
    train_dataset=data["train"],
    eval_dataset=data["valid"],
    quantization_config=bnb,
    peft_config=LoraConfig(r=16, lora_alpha=16, target_modules="all-linear"),
    args=SFTConfig(
        output_dir="runs/qlora",
        learning_rate=1e-4,
        num_train_epochs=2,
        per_device_train_batch_size=2,
        gradient_accumulation_steps=8,
        assistant_only_loss=True,
        eval_strategy="steps",
        eval_steps=50,
        logging_steps=10,
    ),
)
trainer.train()
trainer.save_model("adapter-out")
```

Run it:

```bash
python train_qlora.py
```

Alternative: Axolotl (config-driven):

```bash
axolotl fetch examples
axolotl train examples/llama-3/lora-1b.yml
```

> WARNING: `assistant_only_loss=True` needs a chat template that marks assistant turns. TRL patches the template for known families such as Qwen3; for any other model, read the TRL chat-template notes or remove the flag. A wrong template is the most common cause of a tuned model producing nonsense later.

### Step 5: Or train on an Apple silicon Mac with mlx-lm

Outcome: A LoRA adapter in adapters/ and a test-set perplexity figure.

On a Mac, mlx-lm is the simplest route. It trains LoRA by default, and if you point it at a quantised model it trains QLoRA. Your three JSONL files go in one directory named by `--data`, with the names `train.jsonl`, `valid.jsonl` and `test.jsonl`.

The mlx-lm notes list the model families its LoRA training supports (Mistral, Llama, Phi2, Mixtral, Qwen2, Gemma, OLMo, MiniCPM, InternLM2 at the time of reading). If your base model is newer than that list, run `mlx_lm.lora --help` and try a short run before committing a day to it; if the architecture is not supported the command will tell you early.

To cut memory, lower `--batch-size` (default 4), lower `--num-layers` (default 16), add `--grad-checkpoint`, or shorten examples. The same notes report about 250 tokens per second on an M1 Max with 32 GB for a 7B model with batch size 1 and 4 layers on a sample dataset. Use that only as a rough guide: one million training tokens at 250 tokens per second is about 67 minutes. Run a few steps and read your own rate.

Install:

```bash
pip install "mlx-lm[train]"
```

Train:

```bash
mlx_lm.lora \
  --model <your-base-model> \
  --train \
  --data data \
  --iters 600 \
  --mask-prompt \
  --adapter-path adapters
```

Test-set perplexity with the adapter:

```bash
mlx_lm.lora \
  --model <your-base-model> \
  --adapter-path adapters \
  --data data \
  --test
```

Try the adapter on one prompt:

```bash
mlx_lm.generate \
  --model <your-base-model> \
  --adapter-path adapters \
  --prompt "I was charged twice for March."
```

> NOTE: Perplexity on the test file tells you the model fits your text better. It does not tell you the answers are right. You still need the task-level comparison in step 8.

### Step 6: Merge the adapter into the base model

Outcome: A single full model directory (merged/) with your changes baked in.

A LoRA adapter is a small file that sits on top of the base. To convert to GGUF you first need one ordinary model directory, so merge the adapter into the base weights. With PEFT this is `merge_and_unload()`. The PEFT docs are explicit that it is not an in-place operation: you must keep the model it returns.

Load the base in 16-bit for the merge, not in 4-bit, so the merged weights are not rounded twice. This needs enough memory for the full-precision model, which is about 8 GB for 4B parameters, but it can run on CPU if the GPU is too small.

Other tools have their own merge step. Axolotl uses `axolotl merge-lora` with the same YAML and the adapter directory. Unsloth has `save_pretrained_merged` with `save_method="merged_16bit"`. On a Mac, `mlx_lm.fuse --model <your-base-model>` loads adapters from `adapters/` and writes `fused_model/`; its GGUF export is limited to Mistral, Mixtral and Llama-style models in fp16, so for other families run the fused model with mlx-lm or train on an NVIDIA machine for the GGUF route.

Merge with PEFT (merge.py):

```python
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

BASE = "Qwen/Qwen3-4B"

base = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "adapter-out")
model = model.merge_and_unload()  # returns the merged model; not in place

model.save_pretrained("merged")
AutoTokenizer.from_pretrained(BASE).save_pretrained("merged")
```

Run it:

```bash
python merge.py
```

Alternative: Axolotl:

```bash
axolotl merge-lora train_config.yml --lora-model-dir="./outputs/lora-out"
```

Alternative: mlx-lm:

```bash
mlx_lm.fuse --model <your-base-model>
```

### Step 7: Convert to GGUF and quantise

Outcome: A single .gguf file around a quarter the size of the 16-bit model.

GGUF is the file format llama.cpp and Ollama read. Conversion is two phases: convert the merged Hugging Face model to a high-precision GGUF, then quantise that file. Quantising from 16-bit or 32-bit gives better quality than re-quantising something already quantised.

Get llama.cpp from its repository and build it with CMake. Install its Python requirements, run `convert_hf_to_gguf.py` on the merged directory, then run `llama-quantize` with the method name. `Q4_K_M` is the usual first choice. The llama.cpp notes show Llama 3.1 8B at 14.96 GiB in F16 and 4.58 GiB at Q4_K_M, and 7.95 GiB at Q8_0. Their measured perplexity table is how you would pick a different method if Q4_K_M loses too much quality on your task.

The llama.cpp notes warn that some newer models need a newer `transformers` than its requirements file installs, and say `pip install -U transformers` is safe. If conversion fails with an unknown architecture, update llama.cpp before anything else.

Unsloth users can skip the manual route: `model.save_pretrained_gguf("directory", tokenizer, quantization_method="q4_k_m")` does the merge and quantise in one call. Unsloth also warns that the most common cause of gibberish is a different chat template at inference than at training.

Get and build llama.cpp:

```bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release
python3 -m pip install -r requirements.txt
```

Convert the merged model to GGUF (bf16):

```bash
python convert_hf_to_gguf.py ../merged --outfile my-model-bf16.gguf --outtype bf16
```

Quantise to Q4_K_M:

```bash
./build/bin/llama-quantize my-model-bf16.gguf my-model-Q4_K_M.gguf Q4_K_M
```

> TIP: Keep the bf16 GGUF. If Q4_K_M costs you accuracy on the test set, you can requantise from it to Q5_K_M or Q8_0 without repeating the merge.

### Step 8: Run it locally and compare with the base

Outcome: The tuned model answering in Ollama, with a before and after score on the held-out test set.

Ollama does not quantise GGUF files on import, which is why you did it in the previous step. Create a `Modelfile` that points at the file, build the model with `ollama create`, and then run the same `evaluate.py` against it. Compare the number with the baseline from step 3. If the tuned model beats the base by the margin you set in advance, and nothing else you care about got worse, it passes.

If you would rather use llama.cpp directly, `llama-server -m my-model-Q4_K_M.gguf` starts an OpenAI-compatible server; the server notes say it listens on 127.0.0.1:8080 by default. Point the evaluation script at `http://localhost:8080/v1/` to score it the same way.

Also test what you did not train on. Ask general questions, ask things you want it to refuse, and try inputs in other languages if you serve them. Our post on [safety-first fine-tuning](https://www.swfte.com/blog/safety-first-fine-tuning-explained-2026) lists what to check. For a fuller method covering judges, regression gates and an evidence pack, see [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai).

Modelfile:

```dockerfile
FROM ./my-model-Q4_K_M.gguf
```

Build and try the model:

```bash
ollama create my-model
ollama run my-model "I was charged twice for March."
```

Score the tuned model:

```bash
python evaluate.py my-model data/test.jsonl
```

Alternative: llama.cpp server:

```bash
./build/bin/llama-server -m my-model-Q4_K_M.gguf
```

What a passing comparison looks like (numbers are illustrative):

```text
qwen3:4b: 61/100 exact matches
my-model: 91/100 exact matches
```

### Step 9: Version it, record it and plan its retirement

Outcome: A short model record you can hand to a reviewer, and a rule for when to retrain.

A tuned model is a new artefact someone has to own. Write down: the base model and exact revision, its licence and the date you read it, the dataset version and how examples were approved, training settings, the quantisation method, the baseline and tuned scores, and the safety checks you ran. Store the adapter, the merged weights and the GGUF together with a hash of each.

Decide when to retrain: when the task definition changes, when the base model is superseded, or when monitoring shows the score slipping on fresh cases. Decide when to retire it, and who is allowed to say so.

If you need to serve it to more than one person, put it behind a proper server rather than a laptop. [How to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm) covers that. Swfte does not offer managed fine-tuning today; its Model Vault is designed to hold weights you bring, with versions and promotion stages. Details are on the custom models pages linked below.

## When to stop and use RAG instead

Stop after step 4 if the baseline already passes your threshold: a better prompt may be all you need. Stop after step 6 if the tuned model does not beat the base on the held-out set by a margin you decided in advance. Move to retrieval if the failures on the test set are wrong facts rather than wrong format. Move to a larger base model only after you have a clean evaluation, because otherwise you cannot tell whether the bigger model helped.

If the tuned model improves the task score but degrades something else you care about, such as refusing unsafe requests or answering general questions, treat that as a failed run. Fine-tuning can erode safety behaviour even on harmless data; our post on [safety-first fine-tuning](https://www.swfte.com/blog/safety-first-fine-tuning-explained-2026) explains why and what to test.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| CUDA out of memory during training | Batch size, sequence length or model size is too large for your VRAM. | Lower `per_device_train_batch_size` and raise `gradient_accumulation_steps` by the same factor, shorten examples, or use a smaller base model. Confirm 4-bit loading is on. |
| Training loss falls but validation loss rises | Overfitting to a small dataset. | Stop earlier, reduce epochs, add more varied examples, or lower the learning rate. Do not judge on training loss. |
| The tuned model repeats itself or never stops, or answers in gibberish after conversion | The chat template at inference differs from the one used at training (Unsloth documents this as the most common cause), or the end-of-turn token was not aligned. | Use the same chat template when you run it as when you trained. Test the merged Hugging Face model first; if it is fine there, the problem is in conversion or the Modelfile template. |
| convert_hf_to_gguf.py fails with an unsupported architecture or tokenizer error | llama.cpp or `transformers` is older than the model. | Update the llama.cpp checkout and run `pip install -U transformers`, which the llama.cpp notes say is safe. |
| The tuned model scores no better than the base | Too few or inconsistent examples, a task that is really a knowledge problem, or data leakage that made the baseline look worse than it is. | Read the misses. If they are wrong facts, switch to retrieval. If labels disagree, fix the guideline and relabel. Check train and test do not overlap. |
| mlx_lm.lora errors on your base model | The architecture is not supported by the installed mlx-lm version. | Run `pip install -U mlx-lm`, check `mlx_lm.lora --help`, or choose a model from the supported families listed in the mlx-lm notes. |
| The model got better at the task but worse at declining unsafe requests | Fine-tuning can erode safety behaviour even on harmless data. | Run a refusal and over-refusal suite on base and tuned side by side and gate the release on it. Try a lower rank, fewer epochs, or mix in general examples. |

## Verify it worked

- [ ] `ollama run my-model` answers a test prompt without error and stops at the end of its answer.
- [ ] The tuned model beats the base model on the held-out test file by the margin you wrote down before training.
- [ ] Train, valid and test files share no examples, and nobody tuned against the test file.
- [ ] The same safety and refusal checks pass on the tuned model as on the base.
- [ ] The licence of the base model is saved with the date you read it, and the model record lists dataset version, settings and file hashes.
- [ ] You can rebuild the GGUF from the adapter with the commands in this guide.

## Next steps

- [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai): build a bigger evaluation set, gate releases and keep an evidence pack
- [How to run LLMs locally](https://www.swfte.com/how-to-run-llms-locally): get more from Ollama, llama.cpp and LM Studio, and choose quantisation levels
- [How to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm): serve the model to a team on a GPU server instead of one laptop
- [How to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system): give any model, tuned or not, access to documents that change

## FAQ

### Can I train my own LLM on a laptop?

You can adapt a small open-weight model on a laptop, not train one from scratch. With QLoRA or mlx-lm, models in roughly the 0.5B to 4B range are realistic on a 16 GB machine. Training a new foundation model needs data and compute far beyond a laptop.

### How much data do I need to fine-tune a model?

It depends on the task. OpenAI documents a floor of 10 examples and sees gains from 50 to 100 for its hosted service. For a local adapter, plan on a few hundred reviewed examples to start, and judge by the held-out score, not by a number.

### What is the difference between LoRA and QLoRA?

LoRA trains small adapter matrices while the base weights stay frozen. QLoRA does the same with the base model loaded in 4-bit, which cuts memory a lot. Unsloth describes it as saving about 75 per cent of memory against 16-bit.

### Is fine-tuning better than RAG?

They solve different problems. Fine-tuning changes behaviour: format, tone, narrow classification. RAG supplies knowledge that changes. If answers are wrong because the model lacks facts, use retrieval first.

### What is a GGUF file?

GGUF is the model file format read by llama.cpp and Ollama. It holds the weights, usually quantised to 4 to 8 bits, plus tokeniser metadata in one file, so a model can be copied and loaded on a laptop or server.

### Which quantisation should I use?

Start with Q4_K_M. In the llama.cpp measurements for Llama 3.1 8B it is 4.58 GiB against 14.96 GiB for F16. If it hurts your task score, try Q5_K_M or Q8_0 and keep the bf16 GGUF so you can requantise.

### Does Swfte offer managed fine-tuning?

No, not today. Swfte has a Model Vault designed to hold and deploy weights you bring. The training steps in this guide run on your own hardware or rented GPUs, using open-source tools.

## How Swfte can help

You can do every step above without Swfte. If you want somewhere to keep, version and deploy the weights you produce, or a desktop app to chat with local models, these pages describe what exists.

- [Custom models](https://www.swfte.com/platform/custom-models): how Swfte approaches models you own
- [Deploy and serve](https://www.swfte.com/platform/custom-models/deploy-and-serve): serving weights you bring
- [Deploy models](https://www.swfte.com/deploy-models): self-hosting guides for open-weight models
- [Open-source model testing](https://www.swfte.com/open-source-model-testing): how we test open-weight models before recommending them
- [Cortex](https://www.swfte.com/products/cortex): a desktop app that runs local models through Ollama or LM Studio

Swfte does not offer managed fine-tuning yet: there is no Swfte training service to run for you. Availability and plans for hosting custom models: <custom model availability - founder to fill>.

## Sources

- [Hugging Face model API and model cards](https://huggingface.co/Qwen/Qwen3-4B): licence tag, parameter count and gating of the models in the table, read on 2026-10-06
- [TRL SFTTrainer documentation](https://huggingface.co/docs/trl/sft_trainer): dataset formats, SFTTrainer with peft_config and quantization_config, assistant_only_loss, adapter learning rate
- [PEFT LoRA developer guide](https://huggingface.co/docs/peft/main/en/developer_guides/lora): LoraConfig fields and merge_and_unload() not being in place
- [PEFT quantization guide](https://huggingface.co/docs/peft/developer_guides/quantization): QLoRA definition and target_modules all-linear
- [Transformers bitsandbytes guide](https://huggingface.co/docs/transformers/quantization/bitsandbytes): BitsAndBytesConfig 4-bit and nf4 settings, install line
- [mlx-lm LORA.md](https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/LORA.md): mlx_lm.lora, data layout, test, generate, fuse, memory tips, GGUF export limits
- [mlx-lm README](https://github.com/ml-explore/mlx-lm/blob/main/README.md): pip install mlx-lm
- [Unsloth README](https://github.com/unslothai/unsloth/blob/main/README.md): uv pip install unsloth --torch-backend=auto, supported platforms, licence
- [Unsloth fine-tuning guide](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide): QLoRA memory saving, minimum VRAM, epochs, learning rate
- [Unsloth: saving to GGUF](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf): save_pretrained_gguf, save_pretrained_merged, chat template warning
- [Axolotl README](https://github.com/axolotl-ai-cloud/axolotl/blob/main/README.md): requirements (NVIDIA GPU, Ampere or newer for bf16), axolotl fetch examples and train
- [Axolotl getting started](https://docs.axolotl.ai/docs/getting-started.html): adapter: qlora, axolotl merge-lora, inference commands
- [llama.cpp quantize README](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md): convert_hf_to_gguf.py, llama-quantize Q4_K_M, size table for Llama 3.1 8B
- [llama.cpp build guide](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md): cmake -B build and cmake --build build --config Release
- [llama.cpp README](https://github.com/ggml-org/llama.cpp/blob/master/README.md): install options, and the newer llama cli and llama serve wrapper commands
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md): llama-server default address 127.0.0.1:8080, OpenAI-compatible chat completions
- [Ollama README and download page](https://ollama.com/download): install commands for macOS, Linux and Windows
- [Ollama CLI reference](https://github.com/ollama/ollama/blob/main/docs/cli.mdx): ollama run, pull, create, ps
- [Ollama: importing a model](https://github.com/ollama/ollama/blob/main/docs/import.mdx): Modelfile FROM a GGUF, ollama create, no quantisation on import
- [Ollama OpenAI compatibility](https://docs.ollama.com/openai): base_url http://localhost:11434/v1/, api key ignored
- [OpenAI supervised fine-tuning guide](https://developers.openai.com/api/docs/guides/supervised-fine-tuning): minimum of 10 examples and gains from 50 to 100 for the hosted service
- [Ollama model library: qwen3](https://ollama.com/library/qwen3): qwen3:4b tag exists

Last verified against these sources on 2026-10-06.
