← The journal
Guide Summary

Creating Your Own Local Model: LoRA, GGUF, Ollama Summary

The nine steps to adapt a small open-weight model with LoRA and run it locally with GGUF and Ollama.

Swfte Journal / Guide Summary

This is the short version of our guide, how to create your own local model. The guide has the commands, the expected output and the troubleshooting table, each checked against the tool's own documentation and dated. This post gives you the shape of the work so you can decide whether it is worth doing before you open a terminal.

What "your own model" means here

You are not training a model from scratch. That needs data and compute on a different scale and is not what most teams mean. You are taking a small open-weight model, training a small set of extra weights (a LoRA adapter) on your own examples, merging them into the base, compressing the result into a single GGUF file, and running that file on your own machine with Ollama or llama.cpp.

The guide's one-line answer: it changes behaviour and format well, and facts badly. If your problem is that the model does not know your documents, retrieval is the better tool (see how to build a RAG system). If your problem is that it answers in the wrong style, the wrong structure, or does a narrow task unreliably, adaptation is worth trying.

The nine steps

1. Choose a base model you may legally adapt. Read the licence of the exact checkpoint, not the family. Some licences allow commercial use and modification freely, some add conditions, and some restrict redistribution of derivatives. Record the licence text and the date you read it. This is the cheapest step and the only one that can stop the project outright. Our guide to evaluating an open-source LLM covers the licence check in more detail.

2. Prepare and split your data. Examples of input and output in the shape you want the model to produce. The guide suggests a few hundred to a few thousand good examples, and the word "good" matters more than the number. Hold some back, and do not let them leak into training. Remove personal data you have no basis to use. Duplicates and contradictory examples teach the model confusion.

3. Measure the untouched base model first. Run your held-out cases through the base model and score them. This is your baseline. Without it, you cannot say the adapter helped, and teams skip this step more than any other. Sometimes the base model, with a better prompt, is already good enough, and you have saved yourself the rest.

4. Train a QLoRA adapter on an NVIDIA GPU. The guide uses a verified toolchain and shows the commands. QLoRA trains the adapter while the base model is held in a compressed form, which is what lets a single consumer GPU handle a small model. Memory is the constraint, and the guide shows the arithmetic for working out what fits rather than quoting a figure that will be out of date.

5. Or train on an Apple silicon Mac with mlx-lm. The same idea on the other hardware path, using Apple's MLX framework. Unified memory on a Mac makes small models practical to adapt on a laptop. If you have neither a suitable GPU nor a recent Mac, rent a GPU for the training step only and do everything else locally.

6. Merge the adapter into the base model. The adapter is a separate file until you merge it. Merging produces a full set of weights that behaves like the adapted model without needing the adapter at run time.

7. Convert to GGUF and quantise. llama.cpp supplies the conversion script and a quantisation tool. Quantising shrinks the file and the memory it needs, at some cost to quality. The guide explains the common levels and tells you to check the quality loss on your own held-out set, not on someone else's benchmark.

8. Run it locally and compare with the base. Load the GGUF in Ollama with a Modelfile, or run it directly with llama.cpp. Then run the same held-out cases and compare with the baseline from step 3. If the adapted model is not clearly better on your cases, do not ship it.

9. Version it, record it and plan its retirement. Keep the training data version, the base model checkpoint, the licence, the hyperparameters and the evaluation results together. A model you cannot reproduce or explain is a liability, and a regulator or customer may ask. Decide in advance what would make you retire it: a new base model, a drift in the data, a failed regression test.

Hardware, honestly

The guide derives memory needs from model size and precision, and says to check the tool's documentation for current numbers. As a rule of thumb you can check yourself: the weights of a model with N billion parameters at 16-bit precision take about 2N gigabytes, at 8-bit about N gigabytes, and at 4-bit about N/2 gigabytes, before you add the memory for the context and, during training, for gradients and optimiser state. QLoRA reduces the training cost by holding the base weights at 4-bit. Those are arithmetic facts about numeric formats, not benchmarks of any particular tool.

Time is harder to promise. It depends on data size, sequence length and hardware, and we do not quote figures we have not measured. Run a small trial first and extrapolate.

What a good first project looks like

Pick a narrow job with a clear right answer and plenty of examples you already have. Classifying support tickets into your own categories, rewriting text into your house style, extracting fields from a document type you see every day, or turning notes into a fixed report format are all good first projects. Each has examples sitting in your systems, each has a check you can automate, and each fails in ways you can see.

Poor first projects are the opposite: a general assistant that "knows our business", a model meant to answer questions about a document set that changes weekly, or anything where nobody can agree what a correct output looks like. If a human reviewer cannot grade the output, neither can your held-out set.

Plan for three rounds, not one. The first round mostly teaches you what is wrong with your data: inconsistent labels, examples that contradict each other, cases nobody anticipated. Fixing the data usually improves the result more than changing the training settings. The guide's order, with the baseline measured first, is designed so that each round gives you a number to compare with.

What can go wrong

  • Overfitting. The adapter memorises the training examples and does worse on new ones. The held-out set shows this; the training loss does not.
  • Losing general ability. A model adapted hard on one task can get worse at others. If you still need those, test them too.
  • Quantisation damage. A heavily compressed model can fail on exactly the hard cases you care about. Test the file you will ship, not the one you trained.
  • Chat template mismatch. A model trained with one prompt format and served with another behaves oddly. The guide's troubleshooting table covers the symptoms.
  • Leaked data. Personal or confidential data in the training set can come back out of the model. Clean it first, and treat the resulting model with the same care as the data.

When to stop and use something else

  • If the model lacks knowledge, use retrieval. Fine-tuning does not reliably add facts and can make the model state them confidently and wrongly.
  • If a better prompt fixes it, use the better prompt. It costs nothing to maintain.
  • If you cannot get enough good examples, you are not ready to train. Collect and review more first.
  • If the licence does not allow it, pick a different base model.
  • If you need a hosted service, note that Swfte does not offer managed fine-tuning today. You can do every step above without us, and the guide says so. The platform's custom models pages describe what it is designed to do around models you bring, and the model deployment guides cover serving them.

A checklist before you start

  1. I have read the licence of the exact checkpoint.
  2. I have at least a few hundred reviewed examples and a held-out set.
  3. I have scored the base model on the held-out set.
  4. I know my memory budget for training and for running.
  5. I know what result would make me stop.

If you can tick all five, open the full guide and follow the commands. If you cannot, the first three items are where the next day of work should go.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.