Build · Advanced

How to create your own local model

  • Time: About 1 day for a first working model: a few hours for data, then training time that depends on your hardware and example count.
  • Cost: Free software. You pay for your own hardware, electricity and the hours spent preparing and reviewing examples. Renting a GPU is optional.
  • Level: Advanced
On this page
  1. Short answer
  2. Before you start
  3. First decide: tune, or retrieve?
  4. What your hardware can do: honest sizing
  5. 1. Choose a base model you may legally adapt
  6. 2. Prepare and split your data
  7. 3. Measure the untouched base model first
  8. 4. Train a QLoRA adapter on an NVIDIA GPU
  9. 5. Or train on an Apple silicon Mac with mlx-lm
  10. 6. Merge the adapter into the base model
  11. 7. Convert to GGUF and quantise
  12. 8. Run it locally and compare with the base
  13. 9. Version it, record it and plan its retirement
  14. When to stop and use RAG instead
  15. Troubleshooting
  16. Verify it worked
  17. Next steps
  18. FAQ
  19. How Swfte can help
  20. Sources and last verified

Short answer

You create a local model by taking a small open-weight base model, training a LoRA adapter on a few hundred to a few thousand of your own input and output examples, merging it into the base, quantising the result to GGUF and loading it in Ollama or llama.cpp. It changes behaviour and format well, and facts badly: if the problem is missing knowledge, use retrieval instead. Always compare against the untouched base on held-out cases.

The steps at a glance

  1. Choose a base model you may legally adapt
  2. Prepare and split your data
  3. Measure the untouched base model first
  4. Train a QLoRA adapter on an NVIDIA GPU
  5. Or train on an Apple silicon Mac with mlx-lm
  6. Merge the adapter into the base model
  7. Convert to GGUF and quantise
  8. Run it locally and compare with the base
  9. Version it, record it and plan its retirement

Before you start

Who this is for

  • Engineers who want a small model that answers in a fixed format, tone or classification scheme without sending prompts to an API.
  • Teams with a stable, narrow task and a few hundred reviewed examples of correct output.
  • Anyone who wants to understand the whole path from data file to a model running on their own machine before paying for anything.

Probably not for you if

  • Teams whose real problem is that the model does not know their documents. That is a retrieval problem; see how to build a RAG system.
  • Anyone expecting a model that rivals a frontier API on open-ended work. A small adapter-tuned model is for a narrow job.

Prerequisites

  • A machine you control: an Apple silicon Mac with 16 GB or more of unified memory, or a Linux or Windows machine with an NVIDIA GPU (8 GB of VRAM is the practical floor for a 4B model with QLoRA).
  • Python 3.10 or newer, git, and about 30 GB of free disk for the base weights, merged copy, GGUF files and checkpoints.
  • A few hundred examples of real inputs paired with outputs a qualified person has approved, and the right to train on them. Remove personal data first.
  • Comfort with a terminal and reading a Python script. Nothing here needs a research background.
Time
About 1 day for a first working model: a few hours for data, then training time that depends on your hardware and example count.
Cost
Free software. You pay for your own hardware, electricity and the hours spent preparing and reviewing examples. Renting a GPU is optional.
Hardware
Run-only: about 5 GB of memory for an 8B model at Q4_K_M. Train: 8 to 12 GB of VRAM for a 4B model with QLoRA, or a 16 GB Apple silicon Mac for a small model with mlx-lm. See the sizing table.
Skill
Comfortable with Python and the command line; no machine learning background needed.

Estimates are ours, not measurements, and move with your hardware, data and network.

First decide: tune, or retrieve?

Fine-tuning moves a model's habits: output format, tone, which label to pick, how terse to be. It is a poor way to teach facts that change, because the facts end up frozen in the weights and the model states last quarter's figure with full confidence. Our own post on when to fine-tune sets out the same ladder: better prompt, then retrieval, then fine-tune.

Use this table before you spend a day on training. If your answer is in the left column, stop and use the right-hand fix.

What is going wrongLikely fixWhy
The model does not know our policies, products or documents.Retrieval (build a RAG system)The source of truth stays outside the weights, so updating a document updates the answer.
The answer is right but the format, tone or label set drifts.A clearer prompt with worked examples first; tune if drift persists on a held-out set.This is a behaviour problem, which an adapter handles well.
We run millions of near-identical calls and want a small, cheap model.Tune a small model on outputs you have reviewed.A narrow task with high volume is where a small tuned model pays back.
We have fewer than about a hundred reviewed examples.Write the examples into the prompt and collect more data.Too little data teaches the model your noise, not your task.

What your hardware can do: honest sizing

Memory is the limit that decides everything. A model with N billion parameters stored in 16-bit precision needs about 2 bytes per parameter, so roughly 2N gigabytes for the weights alone. Four-bit quantisation brings that to a little over 0.5 bytes per parameter. The llama.cpp quantisation notes give a measured example: Llama 3.1 8B is 14.96 GiB at F16 and 4.58 GiB at Q4_K_M. Training adds optimiser state, activations and the context length on top, so the training column below is a planning estimate, not a measurement.

Planning estimates. Weights-only figures come from simple arithmetic or the llama.cpp notes; training headroom is our estimate.
TaskModel sizeMemory to plan forBasis
Run only, Q4_K_M8BAbout 5 GB plus contextllama.cpp notes: 4.58 GiB for Llama 3.1 8B at Q4_K_M
Run only, Q4_K_M4BAbout 3 GB plus contextArithmetic: 4B parameters x about 0.6 bytes
QLoRA training4B8 to 12 GB of VRAMEstimate: 4-bit weights about 2 to 3 GB, plus adapter state, activations and context
LoRA training in 16-bit4B16 GB or moreArithmetic: weights alone are about 8 GB, plus training state
LoRA training in 16-bit8B24 GB or moreArithmetic: weights alone are about 16 GB, plus training state
Smallest QLoRA runUnder 1BAbout 3 GBUnsloth fine-tuning guide: as little as 3 GB of VRAM with QLoRA
  1. Step 1Choose a base model you may legally adapt

    You end up with: One named base model, its licence recorded, small enough for your machine.

    Pick an instruction-tuned open-weight model in the 0.5B to 4B range for a first run. Smaller models train faster, fit on modest hardware and make mistakes you can see quickly. Read the licence on the exact checkpoint you will use, because your tuned model inherits it.

    The table shows what the Hugging Face model API reported for these checkpoints on 6 October 2026. Gated means you must accept terms before you can download. If a licence is anything other than Apache-2.0 or MIT, read the licence text itself before you build on it, and ask your legal team if the model will be sold or shipped to customers.

    This guide uses Qwen/Qwen3-4B in its examples, which is listed under Apache-2.0. Swap in any model whose architecture your chosen tool supports; check the tool documentation, not the model card, for that.

    Hugging Face modelParametersLicence tagGated
    Qwen/Qwen3-0.6B0.75 billionapache-2.0No
    HuggingFaceTB/SmolLM3-3B3.08 billionapache-2.0No
    microsoft/Phi-4-mini-instruct3.84 billionmitNo
    Qwen/Qwen3-4B4.02 billionapache-2.0No
    Qwen/Qwen3-8B8.19 billionapache-2.0No
    meta-llama/Llama-3.1-8B-Instruct8.03 billionllama3.1 (custom licence)Yes, manual approval

    Checked against: Hugging Face model API and model cards

  2. Step 2Prepare and split your data

    You end up with: Three JSONL files: train, valid and test, with no overlap between them.

    Write each example as one line of JSON in chat format: a list of messages with the final one from the assistant. This is the conversational format both TRL and mlx-lm read, so the same files work on every path in this guide.

    Quality beats quantity. Include the hard cases and the inputs that currently fail, not only easy ones. If two reviewers would label the same input differently, settle the rule in writing before training, or the model learns the disagreement. Strip personal data and anything you have no right to train on.

    Split before you do anything else, and never look at the test file while tuning. The script below shuffles with a fixed seed and writes an 80, 10, 10 split. It is our script, not from any tool documentation. If the same customer, document or conversation appears in several examples, split by that unit instead, otherwise near-duplicates leak from train into test and flatter the result.

    One example per line (data/all.jsonl) · json
    {"messages": [{"role": "system", "content": "Classify the support ticket as billing, bug or other. Reply with one word."}, {"role": "user", "content": "I was charged twice for March."}, {"role": "assistant", "content": "billing"}]}
    Split into train, valid and test (split.py) · python
    import json, random, pathlib
    
    random.seed(0)
    rows = [json.loads(line) for line in open("data/all.jsonl", encoding="utf-8") if line.strip()]
    random.shuffle(rows)
    
    n = len(rows)
    cut_train, cut_valid = int(n * 0.8), int(n * 0.9)
    parts = {"train": rows[:cut_train], "valid": rows[cut_train:cut_valid], "test": rows[cut_valid:]}
    
    out = pathlib.Path("data")
    for name, part in parts.items():
        with open(out / f"{name}.jsonl", "w", encoding="utf-8") as f:
            for row in part:
                f.write(json.dumps(row, ensure_ascii=False) + "\n")
        print(name, len(part))
    Run it · bash
    python split.py

    Example output for 1,000 examples

    train 800
    valid 100
    test 100

    Checked against: mlx-lm LORA.md, TRL SFTTrainer documentation

  3. Step 3Measure the untouched base model first

    You end up with: A baseline score on your test set that the tuned model must beat.

    Without a baseline you cannot say the tuning helped. Install Ollama, pull the same base model, and score it on the test file. On macOS and Linux the install is one line; on Windows use the PowerShell line. Ollama serves an OpenAI-compatible API on port 11434, so a short script can send each test prompt and compare the reply with the reference.

    The script uses exact match, which suits classification and fixed formats. For free text, replace the comparison with a rubric or a human review; the guide how to validate your AI covers that, including the traps of using a model as the judge. If your model prints reasoning text before the answer, strip it before comparing. Keep temperature at zero so reruns are comparable.

    Record the number and the date. If the base model already clears your bar, stop here: a better prompt is cheaper than a tuned model you now have to own.

    Install Ollama (macOS or Linux) · bash
    curl -fsSL https://ollama.com/install.sh | sh
    Install Ollama (Windows PowerShell) · powershell
    irm https://ollama.com/install.ps1 | iex
    Pull the base model · bash
    ollama pull qwen3:4b
    Score a model on the test set (evaluate.py) · python
    import json, sys
    from openai import OpenAI
    
    client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")  # key is required but ignored
    model, path = sys.argv[1], sys.argv[2]
    
    hits = total = 0
    for line in open(path, encoding="utf-8"):
        if not line.strip():
            continue
        messages = json.loads(line)["messages"]
        prompt, gold = messages[:-1], messages[-1]["content"].strip()
        reply = client.chat.completions.create(model=model, messages=prompt, temperature=0)
        answer = reply.choices[0].message.content.strip()
        hits += int(answer == gold)
        total += 1
        if answer != gold:
            print("MISS | gold:", gold[:60], "| got:", answer[:60])
    
    print(f"{model}: {hits}/{total} exact matches")
    Score the base model · bash
    pip install openai
    python evaluate.py qwen3:4b data/test.jsonl

    Checked against: Ollama README and download page, Ollama OpenAI compatibility, Ollama model library: qwen3

  4. Step 4Train a QLoRA adapter on an NVIDIA GPU

    You end up with: A LoRA adapter saved in adapter-out/, trained on your train file and checked against valid.

    QLoRA loads the base model in 4-bit and trains small adapter matrices on top, which is why it fits in far less memory than training in 16-bit. The Hugging Face docs describe it as quantising the model to 4 bits and then training it with LoRA. TRL, PEFT and bitsandbytes are the standard open-source stack. The script below follows the documented TRL pattern: pass a peft_config and a quantization_config to SFTTrainer, use target_modules="all-linear" as the PEFT quantisation guide advises for QLoRA, and use a learning rate near 1e-4, which the TRL docs suggest for adapters.

    Treat the epoch count, batch size and rank as starting points, not truths. Watch the validation loss. If it falls and then climbs while training loss keeps falling, you are overfitting: stop earlier, use fewer epochs or add data. Unsloth's guide recommends one to three epochs.

    If you prefer a faster, lower-memory wrapper, Unsloth installs with uv pip install unsloth --torch-backend=auto and publishes notebooks for specific models. Axolotl is configuration-driven: you write one YAML file and run axolotl train. Both need an NVIDIA GPU for this path; Axolotl asks for Ampere or newer for bf16 and Flash Attention.

    Install the libraries (Python virtual environment) · bash
    pip install --upgrade transformers accelerate bitsandbytes
    pip install trl peft datasets
    Train (train_qlora.py) · python
    import torch
    from datasets import load_dataset
    from peft import LoraConfig
    from transformers import BitsAndBytesConfig
    from trl import SFTConfig, SFTTrainer
    
    BASE = "Qwen/Qwen3-4B"
    
    data = load_dataset("json", data_files={"train": "data/train.jsonl", "valid": "data/valid.jsonl"})
    
    bnb = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16,
    )
    
    trainer = SFTTrainer(
        model=BASE,
        train_dataset=data["train"],
        eval_dataset=data["valid"],
        quantization_config=bnb,
        peft_config=LoraConfig(r=16, lora_alpha=16, target_modules="all-linear"),
        args=SFTConfig(
            output_dir="runs/qlora",
            learning_rate=1e-4,
            num_train_epochs=2,
            per_device_train_batch_size=2,
            gradient_accumulation_steps=8,
            assistant_only_loss=True,
            eval_strategy="steps",
            eval_steps=50,
            logging_steps=10,
        ),
    )
    trainer.train()
    trainer.save_model("adapter-out")
    Run it · bash
    python train_qlora.py
    Alternative: Axolotl (config-driven) · bash
    axolotl fetch examples
    axolotl train examples/llama-3/lora-1b.yml

    Checked against: TRL SFTTrainer documentation, Transformers bitsandbytes guide, PEFT quantization guide, Unsloth fine-tuning guide, Unsloth README, Axolotl README, Axolotl getting started

  5. Step 5Or train on an Apple silicon Mac with mlx-lm

    You end up with: A LoRA adapter in adapters/ and a test-set perplexity figure.

    On a Mac, mlx-lm is the simplest route. It trains LoRA by default, and if you point it at a quantised model it trains QLoRA. Your three JSONL files go in one directory named by --data, with the names train.jsonl, valid.jsonl and test.jsonl.

    The mlx-lm notes list the model families its LoRA training supports (Mistral, Llama, Phi2, Mixtral, Qwen2, Gemma, OLMo, MiniCPM, InternLM2 at the time of reading). If your base model is newer than that list, run mlx_lm.lora --help and try a short run before committing a day to it; if the architecture is not supported the command will tell you early.

    To cut memory, lower --batch-size (default 4), lower --num-layers (default 16), add --grad-checkpoint, or shorten examples. The same notes report about 250 tokens per second on an M1 Max with 32 GB for a 7B model with batch size 1 and 4 layers on a sample dataset. Use that only as a rough guide: one million training tokens at 250 tokens per second is about 67 minutes. Run a few steps and read your own rate.

    Install · bash
    pip install "mlx-lm[train]"
    Train · bash
    mlx_lm.lora \
      --model <your-base-model> \
      --train \
      --data data \
      --iters 600 \
      --mask-prompt \
      --adapter-path adapters
    Test-set perplexity with the adapter · bash
    mlx_lm.lora \
      --model <your-base-model> \
      --adapter-path adapters \
      --data data \
      --test
    Try the adapter on one prompt · bash
    mlx_lm.generate \
      --model <your-base-model> \
      --adapter-path adapters \
      --prompt "I was charged twice for March."

    Checked against: mlx-lm LORA.md, mlx-lm README

  6. Step 6Merge the adapter into the base model

    You end up with: A single full model directory (merged/) with your changes baked in.

    A LoRA adapter is a small file that sits on top of the base. To convert to GGUF you first need one ordinary model directory, so merge the adapter into the base weights. With PEFT this is merge_and_unload(). The PEFT docs are explicit that it is not an in-place operation: you must keep the model it returns.

    Load the base in 16-bit for the merge, not in 4-bit, so the merged weights are not rounded twice. This needs enough memory for the full-precision model, which is about 8 GB for 4B parameters, but it can run on CPU if the GPU is too small.

    Other tools have their own merge step. Axolotl uses axolotl merge-lora with the same YAML and the adapter directory. Unsloth has save_pretrained_merged with save_method="merged_16bit". On a Mac, mlx_lm.fuse --model <your-base-model> loads adapters from adapters/ and writes fused_model/; its GGUF export is limited to Mistral, Mixtral and Llama-style models in fp16, so for other families run the fused model with mlx-lm or train on an NVIDIA machine for the GGUF route.

    Merge with PEFT (merge.py) · python
    import torch
    from peft import PeftModel
    from transformers import AutoModelForCausalLM, AutoTokenizer
    
    BASE = "Qwen/Qwen3-4B"
    
    base = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16)
    model = PeftModel.from_pretrained(base, "adapter-out")
    model = model.merge_and_unload()  # returns the merged model; not in place
    
    model.save_pretrained("merged")
    AutoTokenizer.from_pretrained(BASE).save_pretrained("merged")
    Run it · bash
    python merge.py
    Alternative: Axolotl · bash
    axolotl merge-lora train_config.yml --lora-model-dir="./outputs/lora-out"
    Alternative: mlx-lm · bash
    mlx_lm.fuse --model <your-base-model>

    Checked against: PEFT LoRA developer guide, Axolotl getting started, Unsloth: saving to GGUF, mlx-lm LORA.md

  7. Step 7Convert to GGUF and quantise

    You end up with: A single .gguf file around a quarter the size of the 16-bit model.

    GGUF is the file format llama.cpp and Ollama read. Conversion is two phases: convert the merged Hugging Face model to a high-precision GGUF, then quantise that file. Quantising from 16-bit or 32-bit gives better quality than re-quantising something already quantised.

    Get llama.cpp from its repository and build it with CMake. Install its Python requirements, run convert_hf_to_gguf.py on the merged directory, then run llama-quantize with the method name. Q4_K_M is the usual first choice. The llama.cpp notes show Llama 3.1 8B at 14.96 GiB in F16 and 4.58 GiB at Q4_K_M, and 7.95 GiB at Q8_0. Their measured perplexity table is how you would pick a different method if Q4_K_M loses too much quality on your task.

    The llama.cpp notes warn that some newer models need a newer transformers than its requirements file installs, and say pip install -U transformers is safe. If conversion fails with an unknown architecture, update llama.cpp before anything else.

    Unsloth users can skip the manual route: model.save_pretrained_gguf("directory", tokenizer, quantization_method="q4_k_m") does the merge and quantise in one call. Unsloth also warns that the most common cause of gibberish is a different chat template at inference than at training.

    Get and build llama.cpp · bash
    git clone https://github.com/ggml-org/llama.cpp
    cd llama.cpp
    cmake -B build
    cmake --build build --config Release
    python3 -m pip install -r requirements.txt
    Convert the merged model to GGUF (bf16) · bash
    python convert_hf_to_gguf.py ../merged --outfile my-model-bf16.gguf --outtype bf16
    Quantise to Q4_K_M · bash
    ./build/bin/llama-quantize my-model-bf16.gguf my-model-Q4_K_M.gguf Q4_K_M

    Checked against: llama.cpp quantize README, llama.cpp build guide, Unsloth: saving to GGUF

  8. Step 8Run it locally and compare with the base

    You end up with: The tuned model answering in Ollama, with a before and after score on the held-out test set.

    Ollama does not quantise GGUF files on import, which is why you did it in the previous step. Create a Modelfile that points at the file, build the model with ollama create, and then run the same evaluate.py against it. Compare the number with the baseline from step 3. If the tuned model beats the base by the margin you set in advance, and nothing else you care about got worse, it passes.

    If you would rather use llama.cpp directly, llama-server -m my-model-Q4_K_M.gguf starts an OpenAI-compatible server; the server notes say it listens on 127.0.0.1:8080 by default. Point the evaluation script at http://localhost:8080/v1/ to score it the same way.

    Also test what you did not train on. Ask general questions, ask things you want it to refuse, and try inputs in other languages if you serve them. Our post on safety-first fine-tuning lists what to check. For a fuller method covering judges, regression gates and an evidence pack, see how to validate your AI.

    Modelfile · dockerfile
    FROM ./my-model-Q4_K_M.gguf
    Build and try the model · bash
    ollama create my-model
    ollama run my-model "I was charged twice for March."
    Score the tuned model · bash
    python evaluate.py my-model data/test.jsonl
    Alternative: llama.cpp server · bash
    ./build/bin/llama-server -m my-model-Q4_K_M.gguf

    What a passing comparison looks like (numbers are illustrative)

    qwen3:4b: 61/100 exact matches
    my-model: 91/100 exact matches

    Checked against: Ollama: importing a model, Ollama CLI reference, llama.cpp server README

  9. Step 9Version it, record it and plan its retirement

    You end up with: A short model record you can hand to a reviewer, and a rule for when to retrain.

    A tuned model is a new artefact someone has to own. Write down: the base model and exact revision, its licence and the date you read it, the dataset version and how examples were approved, training settings, the quantisation method, the baseline and tuned scores, and the safety checks you ran. Store the adapter, the merged weights and the GGUF together with a hash of each.

    Decide when to retrain: when the task definition changes, when the base model is superseded, or when monitoring shows the score slipping on fresh cases. Decide when to retire it, and who is allowed to say so.

    If you need to serve it to more than one person, put it behind a proper server rather than a laptop. How to self-host an LLM covers that. Swfte does not offer managed fine-tuning today; its Model Vault is designed to hold weights you bring, with versions and promotion stages. Details are on the custom models pages linked below.

When to stop and use RAG instead

Stop after step 4 if the baseline already passes your threshold: a better prompt may be all you need. Stop after step 6 if the tuned model does not beat the base on the held-out set by a margin you decided in advance. Move to retrieval if the failures on the test set are wrong facts rather than wrong format. Move to a larger base model only after you have a clean evaluation, because otherwise you cannot tell whether the bigger model helped.

If the tuned model improves the task score but degrades something else you care about, such as refusing unsafe requests or answering general questions, treat that as a failed run. Fine-tuning can erode safety behaviour even on harmless data; our post on safety-first fine-tuning explains why and what to test.

Troubleshooting

What you seeLikely causeFix
CUDA out of memory during trainingBatch size, sequence length or model size is too large for your VRAM.Lower per_device_train_batch_size and raise gradient_accumulation_steps by the same factor, shorten examples, or use a smaller base model. Confirm 4-bit loading is on.
Training loss falls but validation loss risesOverfitting to a small dataset.Stop earlier, reduce epochs, add more varied examples, or lower the learning rate. Do not judge on training loss.
The tuned model repeats itself or never stops, or answers in gibberish after conversionThe chat template at inference differs from the one used at training (Unsloth documents this as the most common cause), or the end-of-turn token was not aligned.Use the same chat template when you run it as when you trained. Test the merged Hugging Face model first; if it is fine there, the problem is in conversion or the Modelfile template.
convert_hf_to_gguf.py fails with an unsupported architecture or tokenizer errorllama.cpp or transformers is older than the model.Update the llama.cpp checkout and run pip install -U transformers, which the llama.cpp notes say is safe.
The tuned model scores no better than the baseToo few or inconsistent examples, a task that is really a knowledge problem, or data leakage that made the baseline look worse than it is.Read the misses. If they are wrong facts, switch to retrieval. If labels disagree, fix the guideline and relabel. Check train and test do not overlap.
mlx_lm.lora errors on your base modelThe architecture is not supported by the installed mlx-lm version.Run pip install -U mlx-lm, check mlx_lm.lora --help, or choose a model from the supported families listed in the mlx-lm notes.
The model got better at the task but worse at declining unsafe requestsFine-tuning can erode safety behaviour even on harmless data.Run a refusal and over-refusal suite on base and tuned side by side and gate the release on it. Try a lower rank, fewer epochs, or mix in general examples.

Verify it worked

Next steps

Related guides

Frequently asked questions

Can I train my own LLM on a laptop?

You can adapt a small open-weight model on a laptop, not train one from scratch. With QLoRA or mlx-lm, models in roughly the 0.5B to 4B range are realistic on a 16 GB machine. Training a new foundation model needs data and compute far beyond a laptop.

How much data do I need to fine-tune a model?

It depends on the task. OpenAI documents a floor of 10 examples and sees gains from 50 to 100 for its hosted service. For a local adapter, plan on a few hundred reviewed examples to start, and judge by the held-out score, not by a number.

What is the difference between LoRA and QLoRA?

LoRA trains small adapter matrices while the base weights stay frozen. QLoRA does the same with the base model loaded in 4-bit, which cuts memory a lot. Unsloth describes it as saving about 75 per cent of memory against 16-bit.

Is fine-tuning better than RAG?

They solve different problems. Fine-tuning changes behaviour: format, tone, narrow classification. RAG supplies knowledge that changes. If answers are wrong because the model lacks facts, use retrieval first.

What is a GGUF file?

GGUF is the model file format read by llama.cpp and Ollama. It holds the weights, usually quantised to 4 to 8 bits, plus tokeniser metadata in one file, so a model can be copied and loaded on a laptop or server.

Which quantisation should I use?

Start with Q4_K_M. In the llama.cpp measurements for Llama 3.1 8B it is 4.58 GiB against 14.96 GiB for F16. If it hurts your task score, try Q5_K_M or Q8_0 and keep the bf16 GGUF so you can requantise.

Does Swfte offer managed fine-tuning?

No, not today. Swfte has a Model Vault designed to hold and deploy weights you bring. The training steps in this guide run on your own hardware or rented GPUs, using open-source tools.

How Swfte can help

You can do every step above without Swfte. If you want somewhere to keep, version and deploy the weights you produce, or a desktop app to chat with local models, these pages describe what exists.

Swfte does not offer managed fine-tuning yet: there is no Swfte training service to run for you. Availability and plans for hosting custom models: <custom model availability - founder to fill>.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. Hugging Face model API and model cards: licence tag, parameter count and gating of the models in the table, read on 2026-10-06
  2. TRL SFTTrainer documentation: dataset formats, SFTTrainer with peft_config and quantization_config, assistant_only_loss, adapter learning rate
  3. PEFT LoRA developer guide: LoraConfig fields and merge_and_unload() not being in place
  4. PEFT quantization guide: QLoRA definition and target_modules all-linear
  5. Transformers bitsandbytes guide: BitsAndBytesConfig 4-bit and nf4 settings, install line
  6. mlx-lm LORA.md: mlx_lm.lora, data layout, test, generate, fuse, memory tips, GGUF export limits
  7. mlx-lm README: pip install mlx-lm
  8. Unsloth README: uv pip install unsloth --torch-backend=auto, supported platforms, licence
  9. Unsloth fine-tuning guide: QLoRA memory saving, minimum VRAM, epochs, learning rate
  10. Unsloth: saving to GGUF: save_pretrained_gguf, save_pretrained_merged, chat template warning
  11. Axolotl README: requirements (NVIDIA GPU, Ampere or newer for bf16), axolotl fetch examples and train
  12. Axolotl getting started: adapter: qlora, axolotl merge-lora, inference commands
  13. llama.cpp quantize README: convert_hf_to_gguf.py, llama-quantize Q4_K_M, size table for Llama 3.1 8B
  14. llama.cpp build guide: cmake -B build and cmake --build build --config Release
  15. llama.cpp README: install options, and the newer llama cli and llama serve wrapper commands
  16. llama.cpp server README: llama-server default address 127.0.0.1:8080, OpenAI-compatible chat completions
  17. Ollama README and download page: install commands for macOS, Linux and Windows
  18. Ollama CLI reference: ollama run, pull, create, ps
  19. Ollama: importing a model: Modelfile FROM a GGUF, ollama create, no quantisation on import
  20. Ollama OpenAI compatibility: base_url http://localhost:11434/v1/, api key ignored
  21. OpenAI supervised fine-tuning guide: minimum of 10 examples and gains from 50 to 100 for the hosted service
  22. Ollama model library: qwen3: qwen3:4b tag exists

Topics

  • lora
  • qlora
  • gguf
  • ollama
  • llama.cpp
  • mlx

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-create-your-own-local-model.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.