Selection guide
Best local LLM: pick a model by memory, task, licence and runtime
Choose a local LLM by the constraint that binds you, using only what the model cards and vendor pages state as of 2026-10-07.
There is no single best local LLM, and this page does not rank one. The model you can run depends on your memory, the task, the licence you can accept and the runtime you use. Eight models are listed below, plus one licence example, each cited to its model card or vendor page, with the rows we could not fully verify named openly. Use the tables to shortlist two or three, then test them on your own tasks.
Last verified 2026-10-07. Sources are listed at the end of the page.
Which local LLM fits my memory?
Sizes below are only what the model card or the runtime’s own library page states. Weights are not the whole bill: the context window and concurrent users add KV cache on top, and Google’s own table says it excludes the context window.
| Memory budget | Model and the figure stated | Where the figure comes from |
|---|---|---|
| Small laptop or edge device | Gemma 4 E2B and E4B. Weights with 20% overhead: E2B 2.9 GB at Q4_0, E4B 4.5 GB at Q4_0 (8.9 GB at SFP8). | Google’s Gemma core docs table. Ollama’s library page lists downloads of 4.6 to 7.5 GB for E2B and 6.6 to 9.5 GB for E4B. |
| About 16 GB | gpt-oss-20b: 21B total and 3.6B active parameters. The card says it runs within 16 GB of memory with MXFP4 weights. | Hugging Face model card. Ollama lists a 14 GB download for the 20b tag. |
| Mid-range laptop or workstation | Granite 4.2 in 3B, 8B and 30B dense sizes. Ollama lists 2.2 GB, 5.3 GB and 18 GB downloads. Gemma 4 12B: 6.7 GB at Q4_0 in Google’s table. | Ollama library page and Google’s Gemma core docs. The Granite card states no memory figure. |
| One 24 GB GPU or a 32 GB Mac | Devstral Small 2 (24B). The card says it can run on a single RTX 4090 or a Mac with 32 GB RAM. | Hugging Face model card. Ollama lists a 15 GB download. |
| One 80 GB data-centre GPU | gpt-oss-120b (117B total, 5.1B active), designed to fit a single 80 GB GPU. Nemotron 3.5 Lightning 30B-A3B: the card gives one H100 80GB or A100 80GB for single-GPU deployment. | Hugging Face model cards. |
| Larger workstation or small server | Gemma 4 31B: 17.5 GB at Q4_0, 34.9 GB at SFP8 and 69.9 GB at BF16, weights only with 20% overhead. Qwen3.8-27B: an 18 GB download on Ollama. | Google’s Gemma core docs table. The Qwen card states no memory figure. |
A download size is not a RAM requirement. Plan for the context length you will use as well.
Which local LLM suits my task?
Each task names the models whose card or library page states that capability. A stated capability is not a measured quality, so test it.
| Task | Models to shortlist | What the source states |
|---|---|---|
| General chat and assistants | gpt-oss-20b, Gemma 4, Qwen3.8-27B, Granite 4.2 | Each is a general instruction model with a permissive licence. Test tone, refusals and formatting on your own prompts. |
| Coding | Devstral Small 2, Qwen3.8-27B | The Devstral card positions it for local deployment on one RTX 4090 or a 32 GB Mac. Qwen3.8 lists coding among its capabilities on Ollama. See best AI coding models and best LLM for coding. |
| Reasoning | gpt-oss-20b, Granite 4.2, Qwen3.8-27B, Gemma 4 | gpt-oss lists three reasoning levels (low, medium, high). Granite lists thinking, non-thinking and low-effort modes. Qwen3.8 has thinking on by default and tunable. Gemma 4 has configurable thinking per Ollama. |
| Multilingual | Gemma 4, Granite 4.2, BGE-M3 for retrieval | Gemma 4: 140+ languages in pre-training and 35+ out of the box. Granite 4.2 names 12 tested languages. BGE-M3 states more than 100 working languages. A language list is not a quality measurement. |
| Vision | Gemma 4, Qwen3.8-27B | Gemma 4 takes image input, with video across the family. Qwen3.8 is a vision-language model with image and video understanding. |
| Embeddings for search and RAG | BGE-M3 | Dense, sparse and multi-vector retrieval, 8,192 tokens of input and 1,024 dimensions on the card. Ollama lists a 1.2 GB download. |
Which licence can I accept for a local LLM?
Licence names are quoted from the model cards. Read the licence file of the exact repository you will download, and ask counsel to confirm your use.
| Model | Licence as the card or licence file names it | What to check |
|---|---|---|
| gpt-oss-20b and gpt-oss-120b | Apache 2.0 | Standard permissive terms. |
| Gemma 4 | Apache 2.0 on the card | Ollama’s library page shows no licence, so confirm on the card for the exact build. |
| Qwen3.8-27B | Apache 2.0 | Read the licence on the card of the exact size you pull. |
| Granite 4.2 | Apache 2.0 | Confirmed on the card and on Ollama’s page. |
| Devstral Small 2 | Apache 2.0 | Check that the repository name matches the card you read. |
| BGE-M3 | MIT on the card | Ollama’s page shows no licence for its build. |
| Nemotron 3.5 Lightning | OpenMDW License Agreement, version 1.1 | A custom licence. Read the text rather than assuming it matches Apache 2.0. |
| Mistral Medium 3.5 (128B) | Modified MIT | The licence says you are not authorised to exercise any rights if your company’s global consolidated monthly revenue exceeds 20 million US dollars for the preceding month. It is an example of an open-weight model many companies cannot use. |
Fine-tunes and quantised builds usually inherit the base licence. Verify that on the repository you use.
Which runtime should I use: Ollama, LM Studio, llama.cpp or vLLM?
The runtime decides how many people can use the model at once and how much setup you carry. The model cards above name these runtimes where the card says so.
Ollama
A command-line and REST API runtime, MIT-licensed per its repository. Its README shows `ollama run gemma4` as the simplest start. Good for one developer or a small tool. See the best Ollama models and Ollama alternatives.
LM Studio
A desktop app that, per its docs, runs models through llama.cpp (GGUF) on Mac, Windows and Linux, and MLX on Apple Silicon, and can serve them on local OpenAI-like endpoints.
llama.cpp
MIT-licensed, aimed at inference with minimal setup on a wide range of hardware. It lists 1.5-bit to 8-bit integer quantisation and includes llama-server, an OpenAI-compatible server. It is the engine under several desktop tools.
vLLM
A GPU serving engine for many simultaneous users, with an OpenAI-compatible server. It suits a shared team endpoint, not a laptop. Compare it with SGLang in SGLang vs vLLM.
Which rows could this guide not fully verify?
- Qwen3.8-27B: the model card states no memory figure and no language list. Ollama’s library page shows no licence, so the Apache 2.0 statement rests on the Hugging Face card.
- Gemma 4: the Hugging Face card and Ollama’s page disagree on which sizes accept audio, so audio is not claimed here. Release dates differ between sources and are omitted.
- Granite 4.2 and Qwen3.8: no memory statement on the card. The only sizes given are Ollama download sizes.
- Nemotron 3.5 Lightning: it is not placed in a runtime row for Ollama or LM Studio, because that support was not checked. The card names PyTorch, vLLM and SGLang.
- BGE-M3: the card states no model size. The only size given is Ollama’s 1.2 GB download.
- LM Studio’s own model catalogue was not read. Runtime support above means the model card or library page names that runtime.
How do I test a local LLM on my own tasks?
An afternoon with your own prompts decides more than any chart. Keep every model on the same prompts and the same settings.
1. Collect 30 real prompts
Take them from real work: a support reply, a contract clause, a code diff. Remove personal data before sharing. Include five that you know are hard.
2. Write the pass rule first
For each prompt, decide in advance what a good answer contains. Use a short checklist so two reviewers would agree.
3. Run each shortlisted model at the quantisation you will ship
A smaller quantisation changes behaviour. Test the file you will deploy, not the full-precision weights.
4. Record speed on your hardware
Note time to first token, tokens per second and memory used at the context length you need. Test two simultaneous users if more than one person will share it.
5. Review blind and repeat
Hide model names while reviewing, and run the set twice. A model that passes once and fails the repeat is not ready.
What does a benchmark not tell you?
A published score tells you how a model did on a fixed test, run by a particular party, in a particular setup. It does not tell you how it will do on your documents, in your language, with your system prompt.
Scores that labs report about their own models are not comparable across labs, because the test version, the harness and the reasoning effort often differ. Quantised builds can behave differently from the full-precision model that was scored. Public tests can leak into training data, and a high score on a narrow task says little about long conversations, refusals or tool use.
Speed claims depend on hardware, batch size and context length. Treat every chart as a hypothesis, and your own 30 prompts as the evidence.
Where Swfte fits
This is a selection guide. Swfte is not a model lab and does not publish a benchmark of these models. It also does not rank its own model: Swfte Safety is a design-intent model card with no measured results, so it is not ranked or listed above.
You do not need Swfte to run any model here. Install a runtime and start. If you later want several models behind one OpenAI-compatible endpoint, with routing, fallbacks, budgets and an audit event stream, Connect is built for that, and its free dashboard works with Ollama, vLLM or LM Studio. Cortex is a desktop app that uses local models through Ollama and LM Studio with on-device knowledge bases. For enterprise-scale choices see best self-hosted models for enterprises, and for the method see how to evaluate an open-source LLM.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Hugging Face: openai/gpt-oss-20b. Licence, parameters, 16 GB memory statement and reasoning levels.
- Hugging Face: openai/gpt-oss-120b. Licence, parameters and single 80 GB GPU statement.
- Hugging Face: google/gemma-4-31B-it. Licence, family sizes, languages, modalities and runtimes named.
- Google: Gemma 4 core docs. Memory table by size and precision, with its overhead assumption.
- Hugging Face: Qwen/Qwen3.8-27B. Licence, vision-language scope, context length and runtimes named.
- Hugging Face: ibm-granite/granite-4.2-30b. Licence, sizes, tested languages and thinking modes.
- Hugging Face: Devstral-Small-2-24B-Instruct-2512. Licence, context window and the RTX 4090 or 32 GB Mac statement.
- Hugging Face: NVIDIA Nemotron 3.5 Lightning 30B-A3B. Licence name, single-GPU statement and runtimes named.
- Hugging Face: BAAI/bge-m3. Licence, input length, languages and retrieval modes.
- Mistral Medium 3.5 licence text. Modified MIT revenue clause.
- Ollama library: gemma4. Download sizes and capabilities.
- Ollama library: gpt-oss. Download sizes and memory statement.
- Ollama library: devstral-small-2. Download size.
- Ollama library: qwen3.8. Download size and capabilities.
- Ollama library: granite4.2. Download sizes.
- Ollama library: bge-m3. Download size.
- Ollama repository. What Ollama is, its licence and the run command.
- LM Studio documentation. Runtimes, formats and the local server.
- llama.cpp README. Goal, hardware, quantisation widths, llama-server and licence.
Frequently asked questions
What is the best local LLM?
It depends on your memory, task, licence and runtime, so this page names no single winner. For a 16 GB machine, gpt-oss-20b states that fit. For one 24 GB GPU or a 32 GB Mac, Devstral Small 2 states it for coding. Shortlist two or three and test them on your own prompts.
How much RAM do I need to run a local LLM?
Weights need roughly the parameter count times the bytes per parameter, plus room for the context window and the runtime. Google’s Gemma 4 table lists 4.5 GB for E4B at Q4_0 and 17.5 GB for the 31B model, with a 20% overhead and no context window included. Check each model card for its own statement.
Is Ollama or LM Studio better for a local LLM?
Neither is better in general. Ollama is a command-line and REST runtime that suits scripts and small tools. LM Studio is a desktop app that runs models through llama.cpp, with MLX on Apple Silicon, and can serve a local OpenAI-like endpoint. Choose by whether you want a terminal or a window.
Can I use a local LLM commercially?
Often yes, but it depends on the licence. The models listed here under Apache 2.0 or MIT allow commercial use on their cards, while Nemotron 3.5 Lightning uses the OpenMDW License Agreement and Mistral Medium 3.5 has a revenue ceiling. Read the licence of the exact repository and ask counsel to confirm.
Are local LLMs private?
A model running on your own machine does not send prompts to a provider, which removes that transfer. Privacy still depends on the rest of your setup, such as logs, plugins, synced folders and any hosted features a tool adds. Check what each tool sends before you rely on it.
Which local LLM is best for coding?
This page names Devstral Small 2 and Qwen3.8-27B as local shortlist entries, because their sources state coding use and a size that fits one workstation. It does not rank them. For a wider look, including hosted models and cited results, see our coding model guides and test on your own repository.