Short answer
To run an LLM locally, install Ollama (one command), pull a small open-weight model that fits your memory, and run it: ollama run <model>. Your prompts then stay on your machine, and the same model is available to code at http://localhost:11434/v1/ through an OpenAI-compatible API. LM Studio gives you a desktop app instead, and llama.cpp gives you the engine directly.
The steps at a glance
Before you start
Who this is for
- Developers and analysts who want a private model on their own machine for drafting, summarising or testing prompts.
- Teams exploring open-weight models before deciding whether to self-host on a server.
- Anyone who needs to work offline, or who must keep prompts off third-party services.
Probably not for you if
- Teams that need a shared service for many users. Start here to learn, then follow how to self-host an LLM.
- Anyone expecting a laptop model to match the largest hosted models. Small local models are good at narrow, well-specified work.
Prerequisites
- A computer with at least 8 GB of memory free for the model: unified memory on an Apple silicon Mac, VRAM on an NVIDIA or AMD GPU, or system RAM (slower).
- About 5 to 10 GB of free disk for your first model; more for larger ones.
- Permission to install software on the machine, and an internet connection for the first download. After that you can work offline.
- Time
- About 20 to 30 minutes, most of it downloading the model.
- Cost
- Free. Ollama, LM Studio and llama.cpp are free to download; you pay for the hardware and electricity you already have.
- Hardware
- Ollama recommends 8 GB of VRAM or unified memory for its small Gemma 4 model. LM Studio recommends 16 GB of RAM on macOS and Windows, and 4 GB of dedicated VRAM on Windows.
- Skill
- No machine learning knowledge needed. You use a terminal for Ollama and llama.cpp; LM Studio has a graphical app.
Estimates are ours, not measurements, and move with your hardware, data and network.
Which tool: Ollama, LM Studio or llama.cpp?
All three run open-weight models on your own machine. They differ in how you drive them, not in what the model can do. Pick one, get a model answering, then try a second if you want to compare.
| Tool | Best for | How you use it | Local API address |
|---|---|---|---|
| Ollama | Developers who want one command and a local API | Terminal and background service | http://localhost:11434 (OpenAI-compatible at /v1/) |
| LM Studio | Anyone who prefers a desktop app, or wants to browse models visually | Graphical app, plus the lms command line | http://localhost:1234/v1 |
| llama.cpp | People who want the engine itself and full control of flags | Command-line programs: llama cli, llama serve | http://127.0.0.1:8080 by default (OpenAI-compatible) |
Step 1Work out how big a model you can run
You end up with: A target model size, in billions of parameters, that fits your machine.
The model has to fit in memory, with room left for the prompt and the conversation so far. A rough rule: a 4-bit file needs a little over half a byte per parameter, so an 8B model is about 5 GB and a 4B model about 3 GB. The llama.cpp notes put Llama 3.1 8B at 4.58 GiB in Q4_K_M. Add a few gigabytes for context, and leave memory for your operating system and browser.
Vendor guidance gives a floor. Ollama's quickstart says its small Gemma 4 model is about a 7.2 GB download and recommends 8 GB of VRAM, or unified memory on a Mac; with less, Ollama can use system RAM but responses may be slower. LM Studio recommends 16 GB of RAM on macOS and Windows, with Apple silicon on macOS 14 or newer and at least 4 GB of dedicated VRAM on Windows.
On a 16 GB machine, aim for a model of about 8B parameters or less in a Q4 file. On 8 GB, aim for 3B to 4B. If a model is slow or crashes, move down one size before trying anything clever.
Your memory for the model A sensible first target Why 8 GB 3B to 4B parameters, Q4 Roughly 2 to 3 GB of weights plus context, leaving room for the system 16 GB Up to about 8B parameters, Q4 to Q5 4.58 GiB for Llama 3.1 8B at Q4_K_M, per llama.cpp 24 GB or more 8B at Q8, or larger models at Q4 Q8_0 for an 8B model is 7.95 GiB; larger models need proportionally more Checked against: llama.cpp quantize README, Ollama quickstart (docs), LM Studio system requirements
Step 2Install Ollama
You end up with: The
ollamacommand works and a local server is running.Ollama is the quickest start. On macOS and Linux, run the one-line installer. On Windows, run the PowerShell line, or download the installer from ollama.com/download. On macOS and Windows you can instead download the app from the same page.
On Linux, if the server is not already running after install, start it with
ollama serve. The Ollama quickstart says to do this when needed. The server listens on 127.0.0.1 port 11434 by default, which means only programs on your machine can reach it.macOS and Linux · bash curl -fsSL https://ollama.com/install.sh | shWindows (PowerShell) · powershell irm https://ollama.com/install.ps1 | iexLinux only, if the server is not running · bash ollama serveChecked against: Ollama download page, Ollama quickstart (docs), Ollama FAQ
Step 3Download a model and talk to it
You end up with: A model answering your questions in the terminal.
Pull a model, then run it. The Ollama docs use
gemma4andgemma4:e2bin their examples, and the quickstart notes the e2b download is about 7.2 GB. Browse ollama.com/library for others and choose a size from step 1. A tag such asqwen3:4bselects a specific size.ollama runpulls the model first if you do not already have it, then opens a chat. Type a question to chat.ollama pslists models currently loaded;ollama stop <model>unloads one;ollama lslists what you have downloaded andollama rm <model>deletes one to free disk.By default a loaded model stays in memory for 5 minutes after the last request, and the default context window is 4096 tokens. Both are configurable: set
OLLAMA_KEEP_ALIVEandOLLAMA_CONTEXT_LENGTHwhen starting the server. Models are stored under~/.ollama/modelson macOS,/usr/share/ollama/.ollama/modelson Linux andC:\Users\%username%\.ollama\modelson Windows; setOLLAMA_MODELSto move them to a bigger disk.Download a model · bash ollama pull gemma4:e2bChat with it · bash ollama run gemma4:e2bManage models · bash ollama ls ollama ps ollama stop gemma4:e2b ollama rm gemma4:e2bChecked against: Ollama CLI reference, Ollama quickstart (docs), Ollama FAQ
Step 4Call the local model from code
You end up with: A script that gets an answer from your local model through the OpenAI-style API.
Ollama exposes an OpenAI-compatible API under
/v1/, so code written for the OpenAI client library usually works with two changes: the base URL and a dummy key. The Ollama docs show exactly this, noting the key is required by the client but ignored. They list support for a subset of the API: chat completions, completions, models, embeddings and responses.This is the same pattern every other tool in this guide uses, which is why it matters: you can switch between Ollama, LM Studio and llama.cpp by changing one URL, and later switch to a server or a hosted API the same way. Keep the model name and base URL in configuration, not in code.
Python (pip install openai first) · python from openai import OpenAI client = OpenAI( base_url="http://localhost:11434/v1/", api_key="ollama", # required but ignored ) reply = client.chat.completions.create( model="gemma4:e2b", messages=[{"role": "user", "content": "Say this is a test"}], ) print(reply.choices[0].message.content)curl · bash curl -X POST http://localhost:11434/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "gemma4:e2b", "messages": [{"role": "user", "content": "Say this is a test"}]}'Checked against: Ollama OpenAI compatibility
Step 5Try LM Studio if you prefer an app
You end up with: The same kind of model running from a desktop app, with its own local server.
Download LM Studio from lmstudio.ai and install it like any other application. Open it, search for a model in its catalogue, download one and chat in the app. Its system requirements are Apple silicon on macOS 14 or newer, Windows on x64 (AVX2 required) or ARM, or Linux as an AppImage on Ubuntu 20.04 or newer.
LM Studio also ships a command line called
lms. Run LM Studio at least once, then open a terminal and enterlms.lms lslists models on disk,lms pslists models loaded in memory,lms getsearches and downloads models andlms server startstarts the local server. The OpenAI-compatible base URL ishttp://localhost:1234/v1, with chat completions, completions, embeddings, models and responses endpoints. Use the Python example above with that URL and any model name thatlms lsshows.Check the command line after running the app once · bash lms ls lms psStart the local server · bash lms server startChecked against: LM Studio system requirements, LM Studio CLI (lms), LM Studio OpenAI compatibility
Step 6Try llama.cpp for direct control
You end up with: A model running straight from the llama.cpp engine, with an OpenAI-compatible server.
llama.cpp is the engine many local tools are built around, and you can run it yourself. Its README gives a one-line installer for macOS and Linux and one for Windows PowerShell, and also lists Docker, prebuilt binaries and building from source. Once installed,
llama cli -hf <repo>downloads a model from Hugging Face and chats with it, andllama serve -hf <repo>starts an OpenAI-compatible server.The README example uses a small
ggml-org/Qwen3.5-0.8B-GGUFmodel, which is a good smoke test because it is tiny. Its server notes say the server listens on 127.0.0.1 port 8080 by default and that the built-in web page is available at the same address. Point the Python example athttp://127.0.0.1:8080/v1/to use it from code.If you already have a GGUF file, for example one you made following how to create your own local model, use
-m <path>with the server program instead of-hf.Install (macOS and Linux) · bash curl -LsSf https://llama.app/install.sh | shInstall (Windows PowerShell) · powershell irm https://llama.app/install.ps1 | iexChat with a small test model · bash llama cli -hf ggml-org/Qwen3.5-0.8B-GGUFStart an OpenAI-compatible server · bash llama serve -hf ggml-org/Qwen3.5-0.8B-GGUFChecked against: llama.cpp README, llama.cpp server README
Step 7Check that it really stays on your machine
You end up with: Confidence that prompts are not leaving the computer, and that the local server is not exposed.
Ollama says that when you run it locally it does not see your prompts or data. Cloud-hosted models are a separate Ollama feature, so check that the model you chose is a local one. Ollama binds to 127.0.0.1 by default; changing
OLLAMA_HOSTto a public address would let other machines reach it, with no login in front, so do not do that on an untrusted network.To prove nothing leaves, disconnect from the network after the download and ask a question. It should still answer. For stricter environments, see how to build an air-gapped AI environment.
Local does not mean unreviewed. A model you download is third-party software with a licence. Read the licence on the model page before using it for work, and prefer files from the publisher or a source you trust.
Checked against: Ollama FAQ, llama.cpp server README
Quantisation levels in plain words
Models are normally stored with 16 bits per number. Quantisation stores them with fewer bits, so the file is smaller and fits in less memory, at some cost in accuracy. The names you will see, such as Q4_K_M or Q8_0, describe the scheme: the number is roughly the bits per weight, so Q4 is about 4 bits and Q8 about 8. If you are unsure, start with a Q4 file.
The llama.cpp maintainers publish measurements for one model, Llama 3.1 8B, in their quantisation notes. They are useful because they show the trade directly. Sizes are exact figures from that table; the speed order is from the same table, run on a single test setup, so your absolute numbers will differ but the order should hold.
| Format | Size (GiB) | Bits per weight | Text generation speed order |
|---|---|---|---|
| F16 (unquantised) | 14.96 | 16.0 | Slowest |
| Q8_0 | 7.95 | 8.50 | Slower |
| Q6_K | 6.14 | 6.56 | Slower |
| Q5_K_M | 5.33 | 5.70 | Faster |
| Q4_K_M | 4.58 | 4.89 | Fast, the usual starting point |
| Q3_K_M | 3.74 | 4.00 | Fast; quality drops further |
| Q2_K | 2.95 | 3.16 | Fastest in that table; expect visible quality loss |
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| The model answers very slowly, one word at a time | The model does not fit in GPU or unified memory, so part of it runs from system RAM or on the CPU. The Ollama quickstart notes this is slower. | Choose a smaller model or a lower quantisation such as Q4_K_M, close other heavy apps, and shorten your prompt. |
| The program crashes or the system freezes when the model loads | The model plus its context needs more memory than you have. | Move down one model size, or reduce the context window. The default for Ollama is 4096 tokens, so check whether you raised it. |
| Code cannot connect to localhost:11434 | The Ollama server is not running, or your code points at the wrong port. | On Linux run ollama serve. Use http://localhost:11434/v1/ for the OpenAI-style API. LM Studio uses port 1234 and llama.cpp 8080. |
| The model forgets the start of a long document | The prompt is longer than the context window. Ollama defaults to 4096 tokens. | Raise OLLAMA_CONTEXT_LENGTH (for example to 8192), or send num_ctx in the request. A bigger context needs more memory. |
| The first answer after a pause takes much longer | The model was unloaded from memory. Ollama keeps models loaded for 5 minutes by default. | Set OLLAMA_KEEP_ALIVE to a longer duration such as 24h if you have the memory to spare. |
| The disk filled up | Each model is several gigabytes, and old ones stay until removed. | Run ollama ls and ollama rm <model> for ones you do not use, or point OLLAMA_MODELS at a larger drive. |
| lms says it is not found | LM Studio has not been run yet, so the command line tool is not set up. | Open LM Studio at least once, then open a new terminal and run lms. |
Verify it worked
Next steps
- How to choose an LLM for your company: turn a good local experiment into a decision with a scorecard
- How to self-host an LLM: serve a model to a team from a GPU server
- How to create your own local model: adapt a small model to your own examples, then run it with these tools
- How to self-host a ChatGPT alternative: put a chat interface in front of Ollama for non-technical colleagues
Related guides
- How to Self-Host an LLM with vLLM (2026 Guide): Serve an open-weight model as a private, OpenAI-compatible endpoint on your own GPU server, with memory sizing, authentication, TLS, metrics and an upgrade routine.
- How to Create Your Own Local Model: LoRA to GGUF: Adapt a small open-weight model to your own examples on one machine, convert it to GGUF, run it locally and check it beats the base model on cases it has not seen.
- How to Choose an LLM for Your Company: A Scorecard: A selection process, not a leaderboard: requirements, a hosted and open-weight shortlist, a test on your own tasks, a weighted scorecard, licence and data-terms checks, and an exit plan.
- How to Self-Host a ChatGPT Alternative (Open WebUI): A hands-on setup of Open WebUI in front of a model you host, with admin accounts, roles, HTTPS, backups and a clear view of where prompts go.
- How to Build an Air-Gapped AI Environment (2026): Stage model weights, container images and Python packages on a connected machine, verify and carry them across, run the model with every online lookup switched off, and prove nothing leaves.
Frequently asked questions
Can I run an LLM locally without a GPU?
Yes, but it is slower. Ollama can use system RAM when VRAM is short, and llama.cpp supports CPU inference. Use a small, quantised model, such as a 3B to 4B model in Q4, and expect slower answers than on a GPU or an Apple silicon Mac.
How much RAM do I need to run a local LLM?
Allow roughly half a byte per parameter for a 4-bit model plus a few gigabytes for context. An 8B model is about 5 GB at Q4_K_M. LM Studio recommends 16 GB of RAM on macOS and Windows; Ollama recommends 8 GB of VRAM or unified memory for its small Gemma 4 model.
Is Ollama or LM Studio better?
They suit different habits. Ollama is a terminal tool with a local API, good for developers. LM Studio is a desktop app with a model browser and a command line, good if you prefer a graphical interface. Both expose an OpenAI-compatible endpoint, so you can try both.
What does Q4_K_M mean?
It is a quantisation format that stores model weights at about 4.9 bits each. In llama.cpp's measurements an 8B model shrinks from 14.96 GiB at F16 to 4.58 GiB at Q4_K_M. It is a common balance of size, speed and quality.
Do local LLMs send my data anywhere?
Ollama states that when you run it locally it does not see your prompts or data. It binds to 127.0.0.1 by default. Confirm by disconnecting from the network after the download, and check that the model you chose is a local model, not a cloud one.
How do I use a local model from my own code?
Point an OpenAI client at the local server: http://localhost:11434/v1/ for Ollama, http://localhost:1234/v1 for LM Studio or http://127.0.0.1:8080 for llama.cpp, with any placeholder API key. Keep the URL and model name in configuration so you can swap them.
Can I run an LLM offline?
Yes. Download the model once while online, then disconnect. All three tools run without a network connection after that. For environments that must never connect, see the air-gapped guide linked in the steps.
How Swfte can help
You can complete this whole guide without Swfte. If you want a desktop app that chats with local models and works with your own documents, or want to compare local and hosted models behind one API, these pages describe what exists.
- Cortex: a desktop app that uses local models through Ollama or LM Studio, with knowledge bases on the device
- Deploy models: guides for moving from a laptop to a server or a private deployment
- Open-source model testing: how we test open-weight models before recommending any
- Connect: one OpenAI-compatible gateway in front of local and hosted models
Everything above works with free tools. Swfte is optional, and nothing here depends on an account.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- Ollama download page: install commands for macOS, Linux and Windows
- Ollama quickstart (docs): gemma4:e2b pull and run, 7.2 GB download, 8 GB VRAM recommendation, ollama serve on Linux
- Ollama CLI reference: ollama run, pull, ls, rm, ps, stop
- Ollama FAQ: default bind address and port, OLLAMA_HOST, OLLAMA_MODELS, keep-alive of 5 minutes, 4096-token default context, local data statement
- Ollama OpenAI compatibility: base_url http://localhost:11434/v1/, ignored api key, supported endpoints
- LM Studio system requirements: macOS, Windows and Linux requirements and recommended RAM and VRAM
- LM Studio CLI (lms): lms get, load, ls, ps and server start; run the app once first
- LM Studio OpenAI compatibility: base URL http://localhost:1234/v1 and supported endpoints
- llama.cpp README: install options, llama cli -hf and llama serve -hf
- llama.cpp server README: default listen address 127.0.0.1:8080
- llama.cpp quantize README: size and bits-per-weight table for Llama 3.1 8B across quantisation formats
Topics
- ollama
- lm studio
- llama.cpp
- quantisation
- local api
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-run-llms-locally.