Selection guide

Best local LLM: pick a model by memory, task, licence and runtime

Choose a local LLM by the constraint that binds you, using only what the model cards and vendor pages state as of 2026-10-07.

There is no single best local LLM, and this page does not rank one. The model you can run depends on your memory, the task, the licence you can accept and the runtime you use. Eight models are listed below, plus one licence example, each cited to its model card or vendor page, with the rows we could not fully verify named openly. Use the tables to shortlist two or three, then test them on your own tasks.

Last verified 2026-10-07. Sources are listed at the end of the page.

Which local LLM fits my memory?

Sizes below are only what the model card or the runtime’s own library page states. Weights are not the whole bill: the context window and concurrent users add KV cache on top, and Google’s own table says it excludes the context window.

Memory budgetModel and the figure statedWhere the figure comes from
Small laptop or edge deviceGemma 4 E2B and E4B. Weights with 20% overhead: E2B 2.9 GB at Q4_0, E4B 4.5 GB at Q4_0 (8.9 GB at SFP8).Google’s Gemma core docs table. Ollama’s library page lists downloads of 4.6 to 7.5 GB for E2B and 6.6 to 9.5 GB for E4B.
About 16 GBgpt-oss-20b: 21B total and 3.6B active parameters. The card says it runs within 16 GB of memory with MXFP4 weights.Hugging Face model card. Ollama lists a 14 GB download for the 20b tag.
Mid-range laptop or workstationGranite 4.2 in 3B, 8B and 30B dense sizes. Ollama lists 2.2 GB, 5.3 GB and 18 GB downloads. Gemma 4 12B: 6.7 GB at Q4_0 in Google’s table.Ollama library page and Google’s Gemma core docs. The Granite card states no memory figure.
One 24 GB GPU or a 32 GB MacDevstral Small 2 (24B). The card says it can run on a single RTX 4090 or a Mac with 32 GB RAM.Hugging Face model card. Ollama lists a 15 GB download.
One 80 GB data-centre GPUgpt-oss-120b (117B total, 5.1B active), designed to fit a single 80 GB GPU. Nemotron 3.5 Lightning 30B-A3B: the card gives one H100 80GB or A100 80GB for single-GPU deployment.Hugging Face model cards.
Larger workstation or small serverGemma 4 31B: 17.5 GB at Q4_0, 34.9 GB at SFP8 and 69.9 GB at BF16, weights only with 20% overhead. Qwen3.8-27B: an 18 GB download on Ollama.Google’s Gemma core docs table. The Qwen card states no memory figure.

A download size is not a RAM requirement. Plan for the context length you will use as well.

Which local LLM suits my task?

Each task names the models whose card or library page states that capability. A stated capability is not a measured quality, so test it.

TaskModels to shortlistWhat the source states
General chat and assistantsgpt-oss-20b, Gemma 4, Qwen3.8-27B, Granite 4.2Each is a general instruction model with a permissive licence. Test tone, refusals and formatting on your own prompts.
CodingDevstral Small 2, Qwen3.8-27BThe Devstral card positions it for local deployment on one RTX 4090 or a 32 GB Mac. Qwen3.8 lists coding among its capabilities on Ollama. See best AI coding models and best LLM for coding.
Reasoninggpt-oss-20b, Granite 4.2, Qwen3.8-27B, Gemma 4gpt-oss lists three reasoning levels (low, medium, high). Granite lists thinking, non-thinking and low-effort modes. Qwen3.8 has thinking on by default and tunable. Gemma 4 has configurable thinking per Ollama.
MultilingualGemma 4, Granite 4.2, BGE-M3 for retrievalGemma 4: 140+ languages in pre-training and 35+ out of the box. Granite 4.2 names 12 tested languages. BGE-M3 states more than 100 working languages. A language list is not a quality measurement.
VisionGemma 4, Qwen3.8-27BGemma 4 takes image input, with video across the family. Qwen3.8 is a vision-language model with image and video understanding.
Embeddings for search and RAGBGE-M3Dense, sparse and multi-vector retrieval, 8,192 tokens of input and 1,024 dimensions on the card. Ollama lists a 1.2 GB download.

Which licence can I accept for a local LLM?

Licence names are quoted from the model cards. Read the licence file of the exact repository you will download, and ask counsel to confirm your use.

ModelLicence as the card or licence file names itWhat to check
gpt-oss-20b and gpt-oss-120bApache 2.0Standard permissive terms.
Gemma 4Apache 2.0 on the cardOllama’s library page shows no licence, so confirm on the card for the exact build.
Qwen3.8-27BApache 2.0Read the licence on the card of the exact size you pull.
Granite 4.2Apache 2.0Confirmed on the card and on Ollama’s page.
Devstral Small 2Apache 2.0Check that the repository name matches the card you read.
BGE-M3MIT on the cardOllama’s page shows no licence for its build.
Nemotron 3.5 LightningOpenMDW License Agreement, version 1.1A custom licence. Read the text rather than assuming it matches Apache 2.0.
Mistral Medium 3.5 (128B)Modified MITThe licence says you are not authorised to exercise any rights if your company’s global consolidated monthly revenue exceeds 20 million US dollars for the preceding month. It is an example of an open-weight model many companies cannot use.

Fine-tunes and quantised builds usually inherit the base licence. Verify that on the repository you use.

Which runtime should I use: Ollama, LM Studio, llama.cpp or vLLM?

The runtime decides how many people can use the model at once and how much setup you carry. The model cards above name these runtimes where the card says so.

Ollama

A command-line and REST API runtime, MIT-licensed per its repository. Its README shows `ollama run gemma4` as the simplest start. Good for one developer or a small tool. See the best Ollama models and Ollama alternatives.

LM Studio

A desktop app that, per its docs, runs models through llama.cpp (GGUF) on Mac, Windows and Linux, and MLX on Apple Silicon, and can serve them on local OpenAI-like endpoints.

llama.cpp

MIT-licensed, aimed at inference with minimal setup on a wide range of hardware. It lists 1.5-bit to 8-bit integer quantisation and includes llama-server, an OpenAI-compatible server. It is the engine under several desktop tools.

vLLM

A GPU serving engine for many simultaneous users, with an OpenAI-compatible server. It suits a shared team endpoint, not a laptop. Compare it with SGLang in SGLang vs vLLM.

Which rows could this guide not fully verify?

  • Qwen3.8-27B: the model card states no memory figure and no language list. Ollama’s library page shows no licence, so the Apache 2.0 statement rests on the Hugging Face card.
  • Gemma 4: the Hugging Face card and Ollama’s page disagree on which sizes accept audio, so audio is not claimed here. Release dates differ between sources and are omitted.
  • Granite 4.2 and Qwen3.8: no memory statement on the card. The only sizes given are Ollama download sizes.
  • Nemotron 3.5 Lightning: it is not placed in a runtime row for Ollama or LM Studio, because that support was not checked. The card names PyTorch, vLLM and SGLang.
  • BGE-M3: the card states no model size. The only size given is Ollama’s 1.2 GB download.
  • LM Studio’s own model catalogue was not read. Runtime support above means the model card or library page names that runtime.

How do I test a local LLM on my own tasks?

An afternoon with your own prompts decides more than any chart. Keep every model on the same prompts and the same settings.

  1. 1. Collect 30 real prompts

    Take them from real work: a support reply, a contract clause, a code diff. Remove personal data before sharing. Include five that you know are hard.

  2. 2. Write the pass rule first

    For each prompt, decide in advance what a good answer contains. Use a short checklist so two reviewers would agree.

  3. 3. Run each shortlisted model at the quantisation you will ship

    A smaller quantisation changes behaviour. Test the file you will deploy, not the full-precision weights.

  4. 4. Record speed on your hardware

    Note time to first token, tokens per second and memory used at the context length you need. Test two simultaneous users if more than one person will share it.

  5. 5. Review blind and repeat

    Hide model names while reviewing, and run the set twice. A model that passes once and fails the repeat is not ready.

What does a benchmark not tell you?

A published score tells you how a model did on a fixed test, run by a particular party, in a particular setup. It does not tell you how it will do on your documents, in your language, with your system prompt.

Scores that labs report about their own models are not comparable across labs, because the test version, the harness and the reasoning effort often differ. Quantised builds can behave differently from the full-precision model that was scored. Public tests can leak into training data, and a high score on a narrow task says little about long conversations, refusals or tool use.

Speed claims depend on hardware, batch size and context length. Treat every chart as a hypothesis, and your own 30 prompts as the evidence.

Where Swfte fits

This is a selection guide. Swfte is not a model lab and does not publish a benchmark of these models. It also does not rank its own model: Swfte Safety is a design-intent model card with no measured results, so it is not ranked or listed above.

You do not need Swfte to run any model here. Install a runtime and start. If you later want several models behind one OpenAI-compatible endpoint, with routing, fallbacks, budgets and an audit event stream, Connect is built for that, and its free dashboard works with Ollama, vLLM or LM Studio. Cortex is a desktop app that uses local models through Ollama and LM Studio with on-device knowledge bases. For enterprise-scale choices see best self-hosted models for enterprises, and for the method see how to evaluate an open-source LLM.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

What is the best local LLM?

It depends on your memory, task, licence and runtime, so this page names no single winner. For a 16 GB machine, gpt-oss-20b states that fit. For one 24 GB GPU or a 32 GB Mac, Devstral Small 2 states it for coding. Shortlist two or three and test them on your own prompts.

How much RAM do I need to run a local LLM?

Weights need roughly the parameter count times the bytes per parameter, plus room for the context window and the runtime. Google’s Gemma 4 table lists 4.5 GB for E4B at Q4_0 and 17.5 GB for the 31B model, with a 20% overhead and no context window included. Check each model card for its own statement.

Is Ollama or LM Studio better for a local LLM?

Neither is better in general. Ollama is a command-line and REST runtime that suits scripts and small tools. LM Studio is a desktop app that runs models through llama.cpp, with MLX on Apple Silicon, and can serve a local OpenAI-like endpoint. Choose by whether you want a terminal or a window.

Can I use a local LLM commercially?

Often yes, but it depends on the licence. The models listed here under Apache 2.0 or MIT allow commercial use on their cards, while Nemotron 3.5 Lightning uses the OpenMDW License Agreement and Mistral Medium 3.5 has a revenue ceiling. Read the licence of the exact repository and ask counsel to confirm.

Are local LLMs private?

A model running on your own machine does not send prompts to a provider, which removes that transfer. Privacy still depends on the rest of your setup, such as logs, plugins, synced folders and any hosted features a tool adds. Check what each tool sends before you rely on it.

Which local LLM is best for coding?

This page names Devstral Small 2 and Qwen3.8-27B as local shortlist entries, because their sources state coding use and a size that fits one workstation. It does not rank them. For a wider look, including hosted models and cited results, see our coding model guides and test on your own repository.

Run local models behind one governed gateway when you outgrow one laptop.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.