Model guide
Best Ollama models: pick one by task and memory, not by rank
A pick-by-task guide built from ten model pages on ollama.com/library as they read on 2026-10-07, with sizes, context, tags and licences as shown.
No single Ollama model is best for every job, and this page does not rank them. It covers ten models that the Ollama library listed on 2026-10-07 and says which task and memory budget each one fits, using only what its page states: the sizes offered, the context window, the capability tags and the licence. Choose by task first, then by the largest size that fits your memory with room left for context, and check the licence before you build on it.
Last verified 2026-10-07. Sources are listed at the end of the page.
Which Ollama model fits which task?
Task fit is paraphrased from the description on each library page. Tags and licence are as shown on the page or its tag page on 2026-10-07.
| Model | What its library page says it is for | Capability tags shown | Licence text shown |
|---|---|---|---|
| gemma4 | Reasoning, agentic workflows, coding and multimodal understanding. Sizes e2b, e4b, 12b, 26b and 31b. | vision, tools, thinking, audio, cloud | Apache License 2.0. Google's model card also says Apache 2.0. |
| qwen3.5 | A family of open-source multimodal models, from 0.8b to 122b, so one family covers a wide range of memory budgets. | vision, tools, thinking | Apache License 2.0. |
| gpt-oss | OpenAI's open-weight models for reasoning, agentic tasks and developer use cases, in 20b and 120b. | tools, thinking, cloud | Apache License 2.0. The page calls it permissive. |
| llama3.2 | Small text models in 1b and 3b sizes, suited to modest hardware. | tools | Llama 3.2 Community Licence Agreement, with an Acceptable Use Policy. |
| granite4 | IBM Granite 4: improved instruction following and tool calling, aimed at enterprise applications. Sizes 350m, 1b and 3b headline the page, with larger "h" variants listed under tags. | tools | Apache License 2.0. |
| deepseek-r1 | Open reasoning models, from 1.5b to 671b. | tools, thinking | MIT licence text. The page adds that distilled models derive from other base models, so read the note for the size you pull. |
| qwen3-coder | Long-context models for agentic and coding tasks, in 30b and 480b. | tools | Apache License 2.0. |
| mistral-small3.2 | A 24b update to Mistral Small with better function calling, instruction following and fewer repetition errors. | vision, tools | Apache License 2.0. |
| embeddinggemma | A 300M-parameter embedding model from Google, for search and retrieval rather than chat. | embedding | Gemma Terms of Use. |
| nomic-embed-text | An open embedding model described as having a large token context window. | embedding | Apache License 2.0. |
How big is each model, and how much context does it offer?
Sizes are the download sizes shown in the library tag tables. A size with a range spans the formats and quantisations listed for that tag. Context is the figure shown for the tag.
| Tag | Size shown | Context shown | Input shown |
|---|---|---|---|
| gemma4:e4b | 6.6GB to 9.5GB | 128K | Text, Image |
| gemma4:26b | 16GB to 19GB | 256K | Text, Image |
| gemma4:31b | 19GB to 20GB | 256K | Text, Image |
| qwen3.5:9b | 6.6GB to 7.6GB | 256K | Text, Image |
| qwen3.5:27b | 17GB to 20GB | 256K | Text, Image |
| gpt-oss:20b | 14GB | 128K | Text |
| gpt-oss:120b | 65GB | 128K | Text |
| llama3.2:3b | 2.0GB | 128K | Text |
| granite4:3b | 2.1GB | 128K | Text |
| granite4:3b-h | 1.9GB | 1M | Text |
| deepseek-r1:8b | 5.2GB | 128K | Text |
| deepseek-r1:32b | 20GB | 128K | Text |
| qwen3-coder:30b | 19GB | 256K | Text |
| mistral-small3.2:24b | 15GB | 128K | Text, Image |
| embeddinggemma:300m | 622MB | 2K | Text |
| nomic-embed-text | 274MB | 2K | Text |
Source: the tag tables on each ollama.com/library page, read on 2026-10-07. The library changes often, so confirm a size before you plan a purchase around it.
How do you read an Ollama library entry?
Start with the capability tags under the model name. Tools means the model is tagged for tool calling, vision means it accepts images, thinking means it has a reasoning mode, embedding means it returns vectors rather than text, audio means audio features, and cloud means a tag runs on Ollama's servers instead of your machine. A tag describes a family, not every size in it, so check the tag table for the size you will pull.
The tag table gives a name such as gemma4:e4b, a download size, a context window and the input types. The name after the colon picks the size, and the default tag, latest, points at one of them. Pull counts and update ages on the library page measure popularity and recency, not quality.
A tag ending in cloud, such as gemma4:31b-cloud, is not local. Your prompts go to Ollama's servers. See Ollama Cloud pricing for what that means.
What do the quantisation tags change?
Quantisation stores weights in fewer bits. A smaller file needs less memory, and Google notes that models with fewer bits are generally less capable than higher-precision ones. These are the tags the library lists for two Gemma 4 sizes.
| Tag | Size shown | What the label tells you |
|---|---|---|
| gemma4:12b-it-qat | 7.2GB | A quantisation-aware training build. Google says QAT builds learn to compensate for the precision loss. |
| gemma4:12b-it-q4_K_M | 8.0GB | A 4-bit quantisation (q4_K_M). |
| gemma4:12b-it-q8_0 | 13GB | An 8-bit quantisation (q8_0). |
| gemma4:12b-it-bf16 | 24GB | The 16-bit weights (bf16). |
| gemma4:31b-it-qat | 19GB | A quantisation-aware training build of the 31B model. |
| gemma4:31b-it-q4_K_M | 20GB | A 4-bit quantisation (q4_K_M). |
| gemma4:31b-it-q8_0 | 34GB | An 8-bit quantisation (q8_0). |
| gemma4:31b-it-bf16 | 63GB | The 16-bit weights (bf16). |
Each size also has MLX tags such as mxfp8 and nvfp4, listed with their own sizes. Source: ollama.com/library/gemma4/tags and ai.google.dev/gemma/docs/core.
How much memory does a model need?
Ollama's own documentation gives no formula for weights and context together, so this page states none of its own. What the sources do say is this. The tag size is the download. Google's Gemma 4 table gives approximate memory to load the weights, including 20 per cent overhead, and says it excludes the context cache: at 4-bit, 17.5 GB for the 31B model and 2.9 GB for E2B. A larger context window needs more on top.
Ollama's context length page says the default context depends on video memory: 4k below 24 GiB, 32k from 24 to 48 GiB, and 256k at 48 GiB or more. It says tasks such as web search, agents and coding tools should use at least 64000 tokens, and that more context needs more memory. The FAQ still says the default is 4096 tokens, so set the value you need explicitly. After loading, run ollama ps and read the PROCESSOR column to see whether the model sits fully on the GPU.
Ollama's FAQ also describes a quantised key and value cache. OLLAMA_KV_CACHE_TYPE can be q8_0, which it says uses about half the memory of f16, or q4_0, about a quarter. It needs Flash Attention, and the effect on quality depends on the model.
How do you choose a model in five steps?
1. Name the task
Chat and writing, coding, image input, tool calling, reasoning, or embeddings for search. Match it to the tags in the task table. For code, also see best AI coding models.
2. Set the memory budget
Take the video memory or unified memory you can give the model, then subtract room for context. Pick the largest tag whose size shown fits with room to spare, and use the Google table for the Gemma family.
3. Check the context you need
Read the context column for the tag and set it explicitly in Ollama. Agents and coding tools need much more than the small default on many machines.
4. Read the licence
Open the licence file on the tag page. See the caveat below, because licences differ even inside one vendor's range.
5. Test on your own prompts
Run ten real tasks through two or three candidates and compare. Use how to evaluate an open-source LLM for the method, and best local LLM for the wider field beyond Ollama's library.
What is the licence caveat?
The Ollama software is MIT-licensed, but every model carries its own licence, and the licences in the task table are not the same. Most of the ten show Apache License 2.0. Llama 3.2 shows a community licence agreement with an acceptable use policy. DeepSeek R1 shows MIT but notes that distilled models derive from other base models. embeddinggemma shows the Gemma Terms of Use, while Gemma 4 shows Apache 2.0, so two models from one family can carry different terms.
Read the licence file on the tag page you pull, rather than the family page alone. Licences can change between versions, and fine-tunes and community uploads carry their own. This page summarises what the pages displayed on 2026-10-07. It is not legal advice, so have counsel read the text before you ship a model inside a product. A cloud tag is also a different service with its own terms.
Where Swfte fits
You do not need Swfte to pull and run an Ollama model, and Swfte does not choose a model for you or rank them. Where several people or services share a runtime, Swfte Connect gives you one OpenAI-compatible API with bring-your-own-key, per-workspace routing rules, budgets and usage caps, content-policy detectors for secrets and personal data, and an audit event stream. These are Built. It lets you change which model a workload uses as a routing decision instead of a setting inside each client. A hosted gateway needs a network path to your Ollama server.
Swfte Cortex is a desktop app that uses local models through Ollama and LM Studio. That is Built. See Connect for the gateway.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Ollama library. Which models the library lists, their descriptions and capability tags.
- Ollama library: gemma4. Sizes, context, input, tags and Apache 2.0 text.
- Ollama library: gemma4 tags. Quantisation and MLX tags with sizes.
- Ollama library: qwen3.5. Sizes, context, tags and licence.
- Ollama library: gpt-oss. Sizes, context, tags and Apache 2.0.
- Ollama library: llama3.2. Sizes, context, tags and community licence.
- Ollama library: granite4. Sizes, context, tags and Apache 2.0.
- Ollama library: deepseek-r1. Sizes, context, tags and MIT note.
- Ollama library: qwen3-coder. Sizes, context, tags and licence.
- Ollama library: mistral-small3.2. Size, context, tags and licence.
- Ollama library: embeddinggemma. Size, context, tags and Gemma Terms of Use.
- Ollama library: nomic-embed-text. Size, context, tags and Apache 2.0.
- Ollama: context length. Default context by video memory and the 64000 token advice.
- Ollama FAQ. Key and value cache quantisation, ollama ps and the 4096 default.
- Ollama quickstart. The 8 GB advice for Gemma 4 E2B.
- Gemma 4 core documentation. Google memory table and the QAT explanation.
Frequently asked questions
What is the best Ollama model?
There is none that is best for every job. The library pages show different strengths: reasoning and agents, coding, image input, small sizes for modest hardware, and embeddings for search. Choose by task, then by the largest size that fits your memory with room for context, then confirm the licence. The task and memory tables above are the starting point.
Which Ollama model should I run with 8 GB of memory?
Ollama's quickstart recommends 8 GB of available video memory or unified memory for Gemma 4 E2B, whose download it puts at about 7.2 GB. Smaller tags in the library tables are llama3.2:3b at 2.0GB, granite4:3b at 2.1GB and deepseek-r1:8b at 5.2GB. Download size is not total memory, because context needs extra.
Which Ollama models can read images?
The library tags gemma4, qwen3.5 and mistral-small3.2 with vision, and their tag tables list Text, Image as the input. Ollama's CLI docs show a Gemma 4 image prompt. Check the tag table for the exact size you pull, because the tag is shown for the family rather than each size.
Which Ollama models are for embeddings?
embeddinggemma and nomic-embed-text carry the embedding tag, and both list a 2K context window. embeddinggemma shows the Gemma Terms of Use and nomic-embed-text shows Apache License 2.0. Embedding models return vectors for retrieval rather than chat answers. See how to build a RAG system for how they are used.
Are Ollama models free for commercial use?
The Ollama software is MIT-licensed, but each model has its own licence, and they differ. Several of the ten here show Apache License 2.0, while Llama 3.2 shows a community licence agreement with an acceptable use policy. Read the licence on the exact tag page before using a model commercially. This is not legal advice.
What does a cloud tag on an Ollama model mean?
A cloud tag, such as gemma4:31b-cloud, runs the model on Ollama's servers instead of your machine, so your prompts leave your device. Ollama says it does not use prompts and responses for training. If you need prompts to stay local, pull a normal tag and consider turning cloud features off.