Deploy · Beginner

How to run LLMs locally

  • Time: About 20 to 30 minutes, most of it downloading the model.
  • Cost: Free. Ollama, LM Studio and llama.cpp are free to download; you pay for the hardware and electricity you already have.
  • Level: Beginner
On this page
  1. Short answer
  2. Before you start
  3. Which tool: Ollama, LM Studio or llama.cpp?
  4. 1. Work out how big a model you can run
  5. 2. Install Ollama
  6. 3. Download a model and talk to it
  7. 4. Call the local model from code
  8. 5. Try LM Studio if you prefer an app
  9. 6. Try llama.cpp for direct control
  10. 7. Check that it really stays on your machine
  11. Quantisation levels in plain words
  12. Troubleshooting
  13. Verify it worked
  14. Next steps
  15. FAQ
  16. How Swfte can help
  17. Sources and last verified

Short answer

To run an LLM locally, install Ollama (one command), pull a small open-weight model that fits your memory, and run it: ollama run <model>. Your prompts then stay on your machine, and the same model is available to code at http://localhost:11434/v1/ through an OpenAI-compatible API. LM Studio gives you a desktop app instead, and llama.cpp gives you the engine directly.

The steps at a glance

  1. Work out how big a model you can run
  2. Install Ollama
  3. Download a model and talk to it
  4. Call the local model from code
  5. Try LM Studio if you prefer an app
  6. Try llama.cpp for direct control
  7. Check that it really stays on your machine

Before you start

Who this is for

  • Developers and analysts who want a private model on their own machine for drafting, summarising or testing prompts.
  • Teams exploring open-weight models before deciding whether to self-host on a server.
  • Anyone who needs to work offline, or who must keep prompts off third-party services.

Probably not for you if

  • Teams that need a shared service for many users. Start here to learn, then follow how to self-host an LLM.
  • Anyone expecting a laptop model to match the largest hosted models. Small local models are good at narrow, well-specified work.

Prerequisites

  • A computer with at least 8 GB of memory free for the model: unified memory on an Apple silicon Mac, VRAM on an NVIDIA or AMD GPU, or system RAM (slower).
  • About 5 to 10 GB of free disk for your first model; more for larger ones.
  • Permission to install software on the machine, and an internet connection for the first download. After that you can work offline.
Time
About 20 to 30 minutes, most of it downloading the model.
Cost
Free. Ollama, LM Studio and llama.cpp are free to download; you pay for the hardware and electricity you already have.
Hardware
Ollama recommends 8 GB of VRAM or unified memory for its small Gemma 4 model. LM Studio recommends 16 GB of RAM on macOS and Windows, and 4 GB of dedicated VRAM on Windows.
Skill
No machine learning knowledge needed. You use a terminal for Ollama and llama.cpp; LM Studio has a graphical app.

Estimates are ours, not measurements, and move with your hardware, data and network.

Which tool: Ollama, LM Studio or llama.cpp?

All three run open-weight models on your own machine. They differ in how you drive them, not in what the model can do. Pick one, get a model answering, then try a second if you want to compare.

ToolBest forHow you use itLocal API address
OllamaDevelopers who want one command and a local APITerminal and background servicehttp://localhost:11434 (OpenAI-compatible at /v1/)
LM StudioAnyone who prefers a desktop app, or wants to browse models visuallyGraphical app, plus the lms command linehttp://localhost:1234/v1
llama.cppPeople who want the engine itself and full control of flagsCommand-line programs: llama cli, llama servehttp://127.0.0.1:8080 by default (OpenAI-compatible)
  1. Step 1Work out how big a model you can run

    You end up with: A target model size, in billions of parameters, that fits your machine.

    The model has to fit in memory, with room left for the prompt and the conversation so far. A rough rule: a 4-bit file needs a little over half a byte per parameter, so an 8B model is about 5 GB and a 4B model about 3 GB. The llama.cpp notes put Llama 3.1 8B at 4.58 GiB in Q4_K_M. Add a few gigabytes for context, and leave memory for your operating system and browser.

    Vendor guidance gives a floor. Ollama's quickstart says its small Gemma 4 model is about a 7.2 GB download and recommends 8 GB of VRAM, or unified memory on a Mac; with less, Ollama can use system RAM but responses may be slower. LM Studio recommends 16 GB of RAM on macOS and Windows, with Apple silicon on macOS 14 or newer and at least 4 GB of dedicated VRAM on Windows.

    On a 16 GB machine, aim for a model of about 8B parameters or less in a Q4 file. On 8 GB, aim for 3B to 4B. If a model is slow or crashes, move down one size before trying anything clever.

    Your memory for the modelA sensible first targetWhy
    8 GB3B to 4B parameters, Q4Roughly 2 to 3 GB of weights plus context, leaving room for the system
    16 GBUp to about 8B parameters, Q4 to Q54.58 GiB for Llama 3.1 8B at Q4_K_M, per llama.cpp
    24 GB or more8B at Q8, or larger models at Q4Q8_0 for an 8B model is 7.95 GiB; larger models need proportionally more

    Checked against: llama.cpp quantize README, Ollama quickstart (docs), LM Studio system requirements

  2. Step 2Install Ollama

    You end up with: The ollama command works and a local server is running.

    Ollama is the quickest start. On macOS and Linux, run the one-line installer. On Windows, run the PowerShell line, or download the installer from ollama.com/download. On macOS and Windows you can instead download the app from the same page.

    On Linux, if the server is not already running after install, start it with ollama serve. The Ollama quickstart says to do this when needed. The server listens on 127.0.0.1 port 11434 by default, which means only programs on your machine can reach it.

    macOS and Linux · bash
    curl -fsSL https://ollama.com/install.sh | sh
    Windows (PowerShell) · powershell
    irm https://ollama.com/install.ps1 | iex
    Linux only, if the server is not running · bash
    ollama serve

    Checked against: Ollama download page, Ollama quickstart (docs), Ollama FAQ

  3. Step 3Download a model and talk to it

    You end up with: A model answering your questions in the terminal.

    Pull a model, then run it. The Ollama docs use gemma4 and gemma4:e2b in their examples, and the quickstart notes the e2b download is about 7.2 GB. Browse ollama.com/library for others and choose a size from step 1. A tag such as qwen3:4b selects a specific size.

    ollama run pulls the model first if you do not already have it, then opens a chat. Type a question to chat. ollama ps lists models currently loaded; ollama stop <model> unloads one; ollama ls lists what you have downloaded and ollama rm <model> deletes one to free disk.

    By default a loaded model stays in memory for 5 minutes after the last request, and the default context window is 4096 tokens. Both are configurable: set OLLAMA_KEEP_ALIVE and OLLAMA_CONTEXT_LENGTH when starting the server. Models are stored under ~/.ollama/models on macOS, /usr/share/ollama/.ollama/models on Linux and C:\Users\%username%\.ollama\models on Windows; set OLLAMA_MODELS to move them to a bigger disk.

    Download a model · bash
    ollama pull gemma4:e2b
    Chat with it · bash
    ollama run gemma4:e2b
    Manage models · bash
    ollama ls
    ollama ps
    ollama stop gemma4:e2b
    ollama rm gemma4:e2b

    Checked against: Ollama CLI reference, Ollama quickstart (docs), Ollama FAQ

  4. Step 4Call the local model from code

    You end up with: A script that gets an answer from your local model through the OpenAI-style API.

    Ollama exposes an OpenAI-compatible API under /v1/, so code written for the OpenAI client library usually works with two changes: the base URL and a dummy key. The Ollama docs show exactly this, noting the key is required by the client but ignored. They list support for a subset of the API: chat completions, completions, models, embeddings and responses.

    This is the same pattern every other tool in this guide uses, which is why it matters: you can switch between Ollama, LM Studio and llama.cpp by changing one URL, and later switch to a server or a hosted API the same way. Keep the model name and base URL in configuration, not in code.

    Python (pip install openai first) · python
    from openai import OpenAI
    
    client = OpenAI(
        base_url="http://localhost:11434/v1/",
        api_key="ollama",  # required but ignored
    )
    
    reply = client.chat.completions.create(
        model="gemma4:e2b",
        messages=[{"role": "user", "content": "Say this is a test"}],
    )
    print(reply.choices[0].message.content)
    curl · bash
    curl -X POST http://localhost:11434/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{"model": "gemma4:e2b", "messages": [{"role": "user", "content": "Say this is a test"}]}'

    Checked against: Ollama OpenAI compatibility

  5. Step 5Try LM Studio if you prefer an app

    You end up with: The same kind of model running from a desktop app, with its own local server.

    Download LM Studio from lmstudio.ai and install it like any other application. Open it, search for a model in its catalogue, download one and chat in the app. Its system requirements are Apple silicon on macOS 14 or newer, Windows on x64 (AVX2 required) or ARM, or Linux as an AppImage on Ubuntu 20.04 or newer.

    LM Studio also ships a command line called lms. Run LM Studio at least once, then open a terminal and enter lms. lms ls lists models on disk, lms ps lists models loaded in memory, lms get searches and downloads models and lms server start starts the local server. The OpenAI-compatible base URL is http://localhost:1234/v1, with chat completions, completions, embeddings, models and responses endpoints. Use the Python example above with that URL and any model name that lms ls shows.

    Check the command line after running the app once · bash
    lms ls
    lms ps
    Start the local server · bash
    lms server start

    Checked against: LM Studio system requirements, LM Studio CLI (lms), LM Studio OpenAI compatibility

  6. Step 6Try llama.cpp for direct control

    You end up with: A model running straight from the llama.cpp engine, with an OpenAI-compatible server.

    llama.cpp is the engine many local tools are built around, and you can run it yourself. Its README gives a one-line installer for macOS and Linux and one for Windows PowerShell, and also lists Docker, prebuilt binaries and building from source. Once installed, llama cli -hf <repo> downloads a model from Hugging Face and chats with it, and llama serve -hf <repo> starts an OpenAI-compatible server.

    The README example uses a small ggml-org/Qwen3.5-0.8B-GGUF model, which is a good smoke test because it is tiny. Its server notes say the server listens on 127.0.0.1 port 8080 by default and that the built-in web page is available at the same address. Point the Python example at http://127.0.0.1:8080/v1/ to use it from code.

    If you already have a GGUF file, for example one you made following how to create your own local model, use -m <path> with the server program instead of -hf.

    Install (macOS and Linux) · bash
    curl -LsSf https://llama.app/install.sh | sh
    Install (Windows PowerShell) · powershell
    irm https://llama.app/install.ps1 | iex
    Chat with a small test model · bash
    llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
    Start an OpenAI-compatible server · bash
    llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

    Checked against: llama.cpp README, llama.cpp server README

  7. Step 7Check that it really stays on your machine

    You end up with: Confidence that prompts are not leaving the computer, and that the local server is not exposed.

    Ollama says that when you run it locally it does not see your prompts or data. Cloud-hosted models are a separate Ollama feature, so check that the model you chose is a local one. Ollama binds to 127.0.0.1 by default; changing OLLAMA_HOST to a public address would let other machines reach it, with no login in front, so do not do that on an untrusted network.

    To prove nothing leaves, disconnect from the network after the download and ask a question. It should still answer. For stricter environments, see how to build an air-gapped AI environment.

    Local does not mean unreviewed. A model you download is third-party software with a licence. Read the licence on the model page before using it for work, and prefer files from the publisher or a source you trust.

    Checked against: Ollama FAQ, llama.cpp server README

Quantisation levels in plain words

Models are normally stored with 16 bits per number. Quantisation stores them with fewer bits, so the file is smaller and fits in less memory, at some cost in accuracy. The names you will see, such as Q4_K_M or Q8_0, describe the scheme: the number is roughly the bits per weight, so Q4 is about 4 bits and Q8 about 8. If you are unsure, start with a Q4 file.

The llama.cpp maintainers publish measurements for one model, Llama 3.1 8B, in their quantisation notes. They are useful because they show the trade directly. Sizes are exact figures from that table; the speed order is from the same table, run on a single test setup, so your absolute numbers will differ but the order should hold.

Llama 3.1 8B in llama.cpp, from its quantisation notes (read 2026-10-06)
FormatSize (GiB)Bits per weightText generation speed order
F16 (unquantised)14.9616.0Slowest
Q8_07.958.50Slower
Q6_K6.146.56Slower
Q5_K_M5.335.70Faster
Q4_K_M4.584.89Fast, the usual starting point
Q3_K_M3.744.00Fast; quality drops further
Q2_K2.953.16Fastest in that table; expect visible quality loss

Troubleshooting

What you seeLikely causeFix
The model answers very slowly, one word at a timeThe model does not fit in GPU or unified memory, so part of it runs from system RAM or on the CPU. The Ollama quickstart notes this is slower.Choose a smaller model or a lower quantisation such as Q4_K_M, close other heavy apps, and shorten your prompt.
The program crashes or the system freezes when the model loadsThe model plus its context needs more memory than you have.Move down one model size, or reduce the context window. The default for Ollama is 4096 tokens, so check whether you raised it.
Code cannot connect to localhost:11434The Ollama server is not running, or your code points at the wrong port.On Linux run ollama serve. Use http://localhost:11434/v1/ for the OpenAI-style API. LM Studio uses port 1234 and llama.cpp 8080.
The model forgets the start of a long documentThe prompt is longer than the context window. Ollama defaults to 4096 tokens.Raise OLLAMA_CONTEXT_LENGTH (for example to 8192), or send num_ctx in the request. A bigger context needs more memory.
The first answer after a pause takes much longerThe model was unloaded from memory. Ollama keeps models loaded for 5 minutes by default.Set OLLAMA_KEEP_ALIVE to a longer duration such as 24h if you have the memory to spare.
The disk filled upEach model is several gigabytes, and old ones stay until removed.Run ollama ls and ollama rm <model> for ones you do not use, or point OLLAMA_MODELS at a larger drive.
lms says it is not foundLM Studio has not been run yet, so the command line tool is not set up.Open LM Studio at least once, then open a new terminal and run lms.

Verify it worked

Next steps

Related guides

Frequently asked questions

Can I run an LLM locally without a GPU?

Yes, but it is slower. Ollama can use system RAM when VRAM is short, and llama.cpp supports CPU inference. Use a small, quantised model, such as a 3B to 4B model in Q4, and expect slower answers than on a GPU or an Apple silicon Mac.

How much RAM do I need to run a local LLM?

Allow roughly half a byte per parameter for a 4-bit model plus a few gigabytes for context. An 8B model is about 5 GB at Q4_K_M. LM Studio recommends 16 GB of RAM on macOS and Windows; Ollama recommends 8 GB of VRAM or unified memory for its small Gemma 4 model.

Is Ollama or LM Studio better?

They suit different habits. Ollama is a terminal tool with a local API, good for developers. LM Studio is a desktop app with a model browser and a command line, good if you prefer a graphical interface. Both expose an OpenAI-compatible endpoint, so you can try both.

What does Q4_K_M mean?

It is a quantisation format that stores model weights at about 4.9 bits each. In llama.cpp's measurements an 8B model shrinks from 14.96 GiB at F16 to 4.58 GiB at Q4_K_M. It is a common balance of size, speed and quality.

Do local LLMs send my data anywhere?

Ollama states that when you run it locally it does not see your prompts or data. It binds to 127.0.0.1 by default. Confirm by disconnecting from the network after the download, and check that the model you chose is a local model, not a cloud one.

How do I use a local model from my own code?

Point an OpenAI client at the local server: http://localhost:11434/v1/ for Ollama, http://localhost:1234/v1 for LM Studio or http://127.0.0.1:8080 for llama.cpp, with any placeholder API key. Keep the URL and model name in configuration so you can swap them.

Can I run an LLM offline?

Yes. Download the model once while online, then disconnect. All three tools run without a network connection after that. For environments that must never connect, see the air-gapped guide linked in the steps.

How Swfte can help

You can complete this whole guide without Swfte. If you want a desktop app that chats with local models and works with your own documents, or want to compare local and hosted models behind one API, these pages describe what exists.

  • Cortex: a desktop app that uses local models through Ollama or LM Studio, with knowledge bases on the device
  • Deploy models: guides for moving from a laptop to a server or a private deployment
  • Open-source model testing: how we test open-weight models before recommending any
  • Connect: one OpenAI-compatible gateway in front of local and hosted models

Everything above works with free tools. Swfte is optional, and nothing here depends on an account.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. Ollama download page: install commands for macOS, Linux and Windows
  2. Ollama quickstart (docs): gemma4:e2b pull and run, 7.2 GB download, 8 GB VRAM recommendation, ollama serve on Linux
  3. Ollama CLI reference: ollama run, pull, ls, rm, ps, stop
  4. Ollama FAQ: default bind address and port, OLLAMA_HOST, OLLAMA_MODELS, keep-alive of 5 minutes, 4096-token default context, local data statement
  5. Ollama OpenAI compatibility: base_url http://localhost:11434/v1/, ignored api key, supported endpoints
  6. LM Studio system requirements: macOS, Windows and Linux requirements and recommended RAM and VRAM
  7. LM Studio CLI (lms): lms get, load, ls, ps and server start; run the app once first
  8. LM Studio OpenAI compatibility: base URL http://localhost:1234/v1 and supported endpoints
  9. llama.cpp README: install options, llama cli -hf and llama serve -hf
  10. llama.cpp server README: default listen address 127.0.0.1:8080
  11. llama.cpp quantize README: size and bits-per-weight table for Llama 3.1 8B across quantisation formats

Topics

  • ollama
  • lm studio
  • llama.cpp
  • quantisation
  • local api

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-run-llms-locally.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.