# How to run LLMs locally

Canonical: https://www.swfte.com/how-to-run-llms-locally
Last verified: 2026-10-06
Difficulty: Beginner
Time: About 20 to 30 minutes, most of it downloading the model.
Cost: Free. Ollama, LM Studio and llama.cpp are free to download; you pay for the hardware and electricity you already have.
Hardware: Ollama recommends 8 GB of VRAM or unified memory for its small Gemma 4 model. LM Studio recommends 16 GB of RAM on macOS and Windows, and 4 GB of dedicated VRAM on Windows.

## Short answer

To run an LLM locally, install Ollama (one command), pull a small open-weight model that fits your memory, and run it: `ollama run <model>`. Your prompts then stay on your machine, and the same model is available to code at http://localhost:11434/v1/ through an OpenAI-compatible API. LM Studio gives you a desktop app instead, and llama.cpp gives you the engine directly.

## Who this is for

- Developers and analysts who want a private model on their own machine for drafting, summarising or testing prompts.
- Teams exploring open-weight models before deciding whether to self-host on a server.
- Anyone who needs to work offline, or who must keep prompts off third-party services.

Not for:
- Teams that need a shared service for many users. Start here to learn, then follow [how to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm).
- Anyone expecting a laptop model to match the largest hosted models. Small local models are good at narrow, well-specified work.

## Prerequisites

- A computer with at least 8 GB of memory free for the model: unified memory on an Apple silicon Mac, VRAM on an NVIDIA or AMD GPU, or system RAM (slower).
- About 5 to 10 GB of free disk for your first model; more for larger ones.
- Permission to install software on the machine, and an internet connection for the first download. After that you can work offline.

## Which tool: Ollama, LM Studio or llama.cpp?

All three run open-weight models on your own machine. They differ in how you drive them, not in what the model can do. Pick one, get a model answering, then try a second if you want to compare.

| Tool | Best for | How you use it | Local API address |
| --- | --- | --- | --- |
| Ollama | Developers who want one command and a local API | Terminal and background service | http://localhost:11434 (OpenAI-compatible at /v1/) |
| LM Studio | Anyone who prefers a desktop app, or wants to browse models visually | Graphical app, plus the `lms` command line | http://localhost:1234/v1 |
| llama.cpp | People who want the engine itself and full control of flags | Command-line programs: `llama cli`, `llama serve` | http://127.0.0.1:8080 by default (OpenAI-compatible) |

> NOTE: More on each: our pages on [Ollama](https://www.swfte.com/tools/ollama), [LM Studio](https://www.swfte.com/tools/lm-studio) and [llama.cpp](https://www.swfte.com/tools/llama-cpp), and the deeper comparison of serving engines in [vLLM vs TGI](https://www.swfte.com/compare/vllm-vs-tgi) if you later move to a server.

## Steps

### Step 1: Work out how big a model you can run

Outcome: A target model size, in billions of parameters, that fits your machine.

The model has to fit in memory, with room left for the prompt and the conversation so far. A rough rule: a 4-bit file needs a little over half a byte per parameter, so an 8B model is about 5 GB and a 4B model about 3 GB. The llama.cpp notes put Llama 3.1 8B at 4.58 GiB in Q4_K_M. Add a few gigabytes for context, and leave memory for your operating system and browser.

Vendor guidance gives a floor. Ollama's quickstart says its small Gemma 4 model is about a 7.2 GB download and recommends 8 GB of VRAM, or unified memory on a Mac; with less, Ollama can use system RAM but responses may be slower. LM Studio recommends 16 GB of RAM on macOS and Windows, with Apple silicon on macOS 14 or newer and at least 4 GB of dedicated VRAM on Windows.

On a 16 GB machine, aim for a model of about 8B parameters or less in a Q4 file. On 8 GB, aim for 3B to 4B. If a model is slow or crashes, move down one size before trying anything clever.

| Your memory for the model | A sensible first target | Why |
| --- | --- | --- |
| 8 GB | 3B to 4B parameters, Q4 | Roughly 2 to 3 GB of weights plus context, leaving room for the system |
| 16 GB | Up to about 8B parameters, Q4 to Q5 | 4.58 GiB for Llama 3.1 8B at Q4_K_M, per llama.cpp |
| 24 GB or more | 8B at Q8, or larger models at Q4 | Q8_0 for an 8B model is 7.95 GiB; larger models need proportionally more |

### Step 2: Install Ollama

Outcome: The `ollama` command works and a local server is running.

Ollama is the quickest start. On macOS and Linux, run the one-line installer. On Windows, run the PowerShell line, or download the installer from ollama.com/download. On macOS and Windows you can instead download the app from the same page.

On Linux, if the server is not already running after install, start it with `ollama serve`. The Ollama quickstart says to do this when needed. The server listens on 127.0.0.1 port 11434 by default, which means only programs on your machine can reach it.

macOS and Linux:

```bash
curl -fsSL https://ollama.com/install.sh | sh
```

Windows (PowerShell):

```powershell
irm https://ollama.com/install.ps1 | iex
```

Linux only, if the server is not running:

```bash
ollama serve
```

> WARNING: Piping a script from the internet into a shell runs it with your permissions. If that is against your policy, download the installer from ollama.com/download and run it normally, or read the script first.

### Step 3: Download a model and talk to it

Outcome: A model answering your questions in the terminal.

Pull a model, then run it. The Ollama docs use `gemma4` and `gemma4:e2b` in their examples, and the quickstart notes the e2b download is about 7.2 GB. Browse ollama.com/library for others and choose a size from step 1. A tag such as `qwen3:4b` selects a specific size.

`ollama run` pulls the model first if you do not already have it, then opens a chat. Type a question to chat. `ollama ps` lists models currently loaded; `ollama stop <model>` unloads one; `ollama ls` lists what you have downloaded and `ollama rm <model>` deletes one to free disk.

By default a loaded model stays in memory for 5 minutes after the last request, and the default context window is 4096 tokens. Both are configurable: set `OLLAMA_KEEP_ALIVE` and `OLLAMA_CONTEXT_LENGTH` when starting the server. Models are stored under `~/.ollama/models` on macOS, `/usr/share/ollama/.ollama/models` on Linux and `C:\Users\%username%\.ollama\models` on Windows; set `OLLAMA_MODELS` to move them to a bigger disk.

Download a model:

```bash
ollama pull gemma4:e2b
```

Chat with it:

```bash
ollama run gemma4:e2b
```

Manage models:

```bash
ollama ls
ollama ps
ollama stop gemma4:e2b
ollama rm gemma4:e2b
```

### Step 4: Call the local model from code

Outcome: A script that gets an answer from your local model through the OpenAI-style API.

Ollama exposes an OpenAI-compatible API under `/v1/`, so code written for the OpenAI client library usually works with two changes: the base URL and a dummy key. The Ollama docs show exactly this, noting the key is required by the client but ignored. They list support for a subset of the API: chat completions, completions, models, embeddings and responses.

This is the same pattern every other tool in this guide uses, which is why it matters: you can switch between Ollama, LM Studio and llama.cpp by changing one URL, and later switch to a server or a hosted API the same way. Keep the model name and base URL in configuration, not in code.

Python (pip install openai first):

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1/",
    api_key="ollama",  # required but ignored
)

reply = client.chat.completions.create(
    model="gemma4:e2b",
    messages=[{"role": "user", "content": "Say this is a test"}],
)
print(reply.choices[0].message.content)
```

curl:

```bash
curl -X POST http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "gemma4:e2b", "messages": [{"role": "user", "content": "Say this is a test"}]}'
```

### Step 5: Try LM Studio if you prefer an app

Outcome: The same kind of model running from a desktop app, with its own local server.

Download LM Studio from lmstudio.ai and install it like any other application. Open it, search for a model in its catalogue, download one and chat in the app. Its system requirements are Apple silicon on macOS 14 or newer, Windows on x64 (AVX2 required) or ARM, or Linux as an AppImage on Ubuntu 20.04 or newer.

LM Studio also ships a command line called `lms`. Run LM Studio at least once, then open a terminal and enter `lms`. `lms ls` lists models on disk, `lms ps` lists models loaded in memory, `lms get` searches and downloads models and `lms server start` starts the local server. The OpenAI-compatible base URL is `http://localhost:1234/v1`, with chat completions, completions, embeddings, models and responses endpoints. Use the Python example above with that URL and any model name that `lms ls` shows.

Check the command line after running the app once:

```bash
lms ls
lms ps
```

Start the local server:

```bash
lms server start
```

### Step 6: Try llama.cpp for direct control

Outcome: A model running straight from the llama.cpp engine, with an OpenAI-compatible server.

llama.cpp is the engine many local tools are built around, and you can run it yourself. Its README gives a one-line installer for macOS and Linux and one for Windows PowerShell, and also lists Docker, prebuilt binaries and building from source. Once installed, `llama cli -hf <repo>` downloads a model from Hugging Face and chats with it, and `llama serve -hf <repo>` starts an OpenAI-compatible server.

The README example uses a small `ggml-org/Qwen3.5-0.8B-GGUF` model, which is a good smoke test because it is tiny. Its server notes say the server listens on 127.0.0.1 port 8080 by default and that the built-in web page is available at the same address. Point the Python example at `http://127.0.0.1:8080/v1/` to use it from code.

If you already have a GGUF file, for example one you made following [how to create your own local model](https://www.swfte.com/how-to-create-your-own-local-model), use `-m <path>` with the server program instead of `-hf`.

Install (macOS and Linux):

```bash
curl -LsSf https://llama.app/install.sh | sh
```

Install (Windows PowerShell):

```powershell
irm https://llama.app/install.ps1 | iex
```

Chat with a small test model:

```bash
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
```

Start an OpenAI-compatible server:

```bash
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
```

### Step 7: Check that it really stays on your machine

Outcome: Confidence that prompts are not leaving the computer, and that the local server is not exposed.

Ollama says that when you run it locally it does not see your prompts or data. Cloud-hosted models are a separate Ollama feature, so check that the model you chose is a local one. Ollama binds to 127.0.0.1 by default; changing `OLLAMA_HOST` to a public address would let other machines reach it, with no login in front, so do not do that on an untrusted network.

To prove nothing leaves, disconnect from the network after the download and ask a question. It should still answer. For stricter environments, see [how to build an air-gapped AI environment](https://www.swfte.com/how-to-build-an-air-gapped-ai-environment).

Local does not mean unreviewed. A model you download is third-party software with a licence. Read the licence on the model page before using it for work, and prefer files from the publisher or a source you trust.

## Quantisation levels in plain words

Models are normally stored with 16 bits per number. Quantisation stores them with fewer bits, so the file is smaller and fits in less memory, at some cost in accuracy. The names you will see, such as Q4_K_M or Q8_0, describe the scheme: the number is roughly the bits per weight, so Q4 is about 4 bits and Q8 about 8. If you are unsure, start with a Q4 file.

The llama.cpp maintainers publish measurements for one model, Llama 3.1 8B, in their quantisation notes. They are useful because they show the trade directly. Sizes are exact figures from that table; the speed order is from the same table, run on a single test setup, so your absolute numbers will differ but the order should hold.

**Llama 3.1 8B in llama.cpp, from its quantisation notes (read 2026-10-06)**

| Format | Size (GiB) | Bits per weight | Text generation speed order |
| --- | --- | --- | --- |
| F16 (unquantised) | 14.96 | 16.0 | Slowest |
| Q8_0 | 7.95 | 8.50 | Slower |
| Q6_K | 6.14 | 6.56 | Slower |
| Q5_K_M | 5.33 | 5.70 | Faster |
| Q4_K_M | 4.58 | 4.89 | Fast, the usual starting point |
| Q3_K_M | 3.74 | 4.00 | Fast; quality drops further |
| Q2_K | 2.95 | 3.16 | Fastest in that table; expect visible quality loss |

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| The model answers very slowly, one word at a time | The model does not fit in GPU or unified memory, so part of it runs from system RAM or on the CPU. The Ollama quickstart notes this is slower. | Choose a smaller model or a lower quantisation such as Q4_K_M, close other heavy apps, and shorten your prompt. |
| The program crashes or the system freezes when the model loads | The model plus its context needs more memory than you have. | Move down one model size, or reduce the context window. The default for Ollama is 4096 tokens, so check whether you raised it. |
| Code cannot connect to localhost:11434 | The Ollama server is not running, or your code points at the wrong port. | On Linux run `ollama serve`. Use http://localhost:11434/v1/ for the OpenAI-style API. LM Studio uses port 1234 and llama.cpp 8080. |
| The model forgets the start of a long document | The prompt is longer than the context window. Ollama defaults to 4096 tokens. | Raise `OLLAMA_CONTEXT_LENGTH` (for example to 8192), or send `num_ctx` in the request. A bigger context needs more memory. |
| The first answer after a pause takes much longer | The model was unloaded from memory. Ollama keeps models loaded for 5 minutes by default. | Set `OLLAMA_KEEP_ALIVE` to a longer duration such as 24h if you have the memory to spare. |
| The disk filled up | Each model is several gigabytes, and old ones stay until removed. | Run `ollama ls` and `ollama rm <model>` for ones you do not use, or point `OLLAMA_MODELS` at a larger drive. |
| lms says it is not found | LM Studio has not been run yet, so the command line tool is not set up. | Open LM Studio at least once, then open a new terminal and run `lms`. |

## Verify it worked

- [ ] `ollama run <model>` answers a plain question and `ollama ps` shows the model loaded.
- [ ] The Python or curl example returns text from http://localhost:11434/v1/.
- [ ] With the network disconnected, the model still answers.
- [ ] You know which quantisation file you downloaded and how much memory it uses.
- [ ] The licence for the model you chose is recorded somewhere you can find it.
- [ ] Your local server listens on 127.0.0.1 only, unless you decided otherwise on purpose.

## Next steps

- [How to choose an LLM for your company](https://www.swfte.com/how-to-choose-an-llm-for-your-company): turn a good local experiment into a decision with a scorecard
- [How to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm): serve a model to a team from a GPU server
- [How to create your own local model](https://www.swfte.com/how-to-create-your-own-local-model): adapt a small model to your own examples, then run it with these tools
- [How to self-host a ChatGPT alternative](https://www.swfte.com/how-to-self-host-chatgpt-alternative): put a chat interface in front of Ollama for non-technical colleagues

## FAQ

### Can I run an LLM locally without a GPU?

Yes, but it is slower. Ollama can use system RAM when VRAM is short, and llama.cpp supports CPU inference. Use a small, quantised model, such as a 3B to 4B model in Q4, and expect slower answers than on a GPU or an Apple silicon Mac.

### How much RAM do I need to run a local LLM?

Allow roughly half a byte per parameter for a 4-bit model plus a few gigabytes for context. An 8B model is about 5 GB at Q4_K_M. LM Studio recommends 16 GB of RAM on macOS and Windows; Ollama recommends 8 GB of VRAM or unified memory for its small Gemma 4 model.

### Is Ollama or LM Studio better?

They suit different habits. Ollama is a terminal tool with a local API, good for developers. LM Studio is a desktop app with a model browser and a command line, good if you prefer a graphical interface. Both expose an OpenAI-compatible endpoint, so you can try both.

### What does Q4_K_M mean?

It is a quantisation format that stores model weights at about 4.9 bits each. In llama.cpp's measurements an 8B model shrinks from 14.96 GiB at F16 to 4.58 GiB at Q4_K_M. It is a common balance of size, speed and quality.

### Do local LLMs send my data anywhere?

Ollama states that when you run it locally it does not see your prompts or data. It binds to 127.0.0.1 by default. Confirm by disconnecting from the network after the download, and check that the model you chose is a local model, not a cloud one.

### How do I use a local model from my own code?

Point an OpenAI client at the local server: http://localhost:11434/v1/ for Ollama, http://localhost:1234/v1 for LM Studio or http://127.0.0.1:8080 for llama.cpp, with any placeholder API key. Keep the URL and model name in configuration so you can swap them.

### Can I run an LLM offline?

Yes. Download the model once while online, then disconnect. All three tools run without a network connection after that. For environments that must never connect, see the air-gapped guide linked in the steps.

## How Swfte can help

You can complete this whole guide without Swfte. If you want a desktop app that chats with local models and works with your own documents, or want to compare local and hosted models behind one API, these pages describe what exists.

- [Cortex](https://www.swfte.com/products/cortex): a desktop app that uses local models through Ollama or LM Studio, with knowledge bases on the device
- [Deploy models](https://www.swfte.com/deploy-models): guides for moving from a laptop to a server or a private deployment
- [Open-source model testing](https://www.swfte.com/open-source-model-testing): how we test open-weight models before recommending any
- [Connect](https://www.swfte.com/products/connect): one OpenAI-compatible gateway in front of local and hosted models

Everything above works with free tools. Swfte is optional, and nothing here depends on an account.

## Sources

- [Ollama download page](https://ollama.com/download): install commands for macOS, Linux and Windows
- [Ollama quickstart (docs)](https://github.com/ollama/ollama/blob/main/docs/quickstart.mdx): gemma4:e2b pull and run, 7.2 GB download, 8 GB VRAM recommendation, ollama serve on Linux
- [Ollama CLI reference](https://github.com/ollama/ollama/blob/main/docs/cli.mdx): ollama run, pull, ls, rm, ps, stop
- [Ollama FAQ](https://docs.ollama.com/faq): default bind address and port, OLLAMA_HOST, OLLAMA_MODELS, keep-alive of 5 minutes, 4096-token default context, local data statement
- [Ollama OpenAI compatibility](https://docs.ollama.com/openai): base_url http://localhost:11434/v1/, ignored api key, supported endpoints
- [LM Studio system requirements](https://lmstudio.ai/docs/app/system-requirements): macOS, Windows and Linux requirements and recommended RAM and VRAM
- [LM Studio CLI (lms)](https://lmstudio.ai/docs/cli): lms get, load, ls, ps and server start; run the app once first
- [LM Studio OpenAI compatibility](https://lmstudio.ai/docs/developer/openai-compat): base URL http://localhost:1234/v1 and supported endpoints
- [llama.cpp README](https://github.com/ggml-org/llama.cpp/blob/master/README.md): install options, llama cli -hf and llama serve -hf
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md): default listen address 127.0.0.1:8080
- [llama.cpp quantize README](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md): size and bits-per-weight table for Llama 3.1 8B across quantisation formats

Last verified against these sources on 2026-10-06.
