Model guide

Best Ollama models: pick one by task and memory, not by rank

A pick-by-task guide built from ten model pages on ollama.com/library as they read on 2026-10-07, with sizes, context, tags and licences as shown.

No single Ollama model is best for every job, and this page does not rank them. It covers ten models that the Ollama library listed on 2026-10-07 and says which task and memory budget each one fits, using only what its page states: the sizes offered, the context window, the capability tags and the licence. Choose by task first, then by the largest size that fits your memory with room left for context, and check the licence before you build on it.

Last verified 2026-10-07. Sources are listed at the end of the page.

Which Ollama model fits which task?

Task fit is paraphrased from the description on each library page. Tags and licence are as shown on the page or its tag page on 2026-10-07.

ModelWhat its library page says it is forCapability tags shownLicence text shown
gemma4Reasoning, agentic workflows, coding and multimodal understanding. Sizes e2b, e4b, 12b, 26b and 31b.vision, tools, thinking, audio, cloudApache License 2.0. Google's model card also says Apache 2.0.
qwen3.5A family of open-source multimodal models, from 0.8b to 122b, so one family covers a wide range of memory budgets.vision, tools, thinkingApache License 2.0.
gpt-ossOpenAI's open-weight models for reasoning, agentic tasks and developer use cases, in 20b and 120b.tools, thinking, cloudApache License 2.0. The page calls it permissive.
llama3.2Small text models in 1b and 3b sizes, suited to modest hardware.toolsLlama 3.2 Community Licence Agreement, with an Acceptable Use Policy.
granite4IBM Granite 4: improved instruction following and tool calling, aimed at enterprise applications. Sizes 350m, 1b and 3b headline the page, with larger "h" variants listed under tags.toolsApache License 2.0.
deepseek-r1Open reasoning models, from 1.5b to 671b.tools, thinkingMIT licence text. The page adds that distilled models derive from other base models, so read the note for the size you pull.
qwen3-coderLong-context models for agentic and coding tasks, in 30b and 480b.toolsApache License 2.0.
mistral-small3.2A 24b update to Mistral Small with better function calling, instruction following and fewer repetition errors.vision, toolsApache License 2.0.
embeddinggemmaA 300M-parameter embedding model from Google, for search and retrieval rather than chat.embeddingGemma Terms of Use.
nomic-embed-textAn open embedding model described as having a large token context window.embeddingApache License 2.0.

How big is each model, and how much context does it offer?

Sizes are the download sizes shown in the library tag tables. A size with a range spans the formats and quantisations listed for that tag. Context is the figure shown for the tag.

TagSize shownContext shownInput shown
gemma4:e4b6.6GB to 9.5GB128KText, Image
gemma4:26b16GB to 19GB256KText, Image
gemma4:31b19GB to 20GB256KText, Image
qwen3.5:9b6.6GB to 7.6GB256KText, Image
qwen3.5:27b17GB to 20GB256KText, Image
gpt-oss:20b14GB128KText
gpt-oss:120b65GB128KText
llama3.2:3b2.0GB128KText
granite4:3b2.1GB128KText
granite4:3b-h1.9GB1MText
deepseek-r1:8b5.2GB128KText
deepseek-r1:32b20GB128KText
qwen3-coder:30b19GB256KText
mistral-small3.2:24b15GB128KText, Image
embeddinggemma:300m622MB2KText
nomic-embed-text274MB2KText

Source: the tag tables on each ollama.com/library page, read on 2026-10-07. The library changes often, so confirm a size before you plan a purchase around it.

How do you read an Ollama library entry?

Start with the capability tags under the model name. Tools means the model is tagged for tool calling, vision means it accepts images, thinking means it has a reasoning mode, embedding means it returns vectors rather than text, audio means audio features, and cloud means a tag runs on Ollama's servers instead of your machine. A tag describes a family, not every size in it, so check the tag table for the size you will pull.

The tag table gives a name such as gemma4:e4b, a download size, a context window and the input types. The name after the colon picks the size, and the default tag, latest, points at one of them. Pull counts and update ages on the library page measure popularity and recency, not quality.

A tag ending in cloud, such as gemma4:31b-cloud, is not local. Your prompts go to Ollama's servers. See Ollama Cloud pricing for what that means.

What do the quantisation tags change?

Quantisation stores weights in fewer bits. A smaller file needs less memory, and Google notes that models with fewer bits are generally less capable than higher-precision ones. These are the tags the library lists for two Gemma 4 sizes.

TagSize shownWhat the label tells you
gemma4:12b-it-qat7.2GBA quantisation-aware training build. Google says QAT builds learn to compensate for the precision loss.
gemma4:12b-it-q4_K_M8.0GBA 4-bit quantisation (q4_K_M).
gemma4:12b-it-q8_013GBAn 8-bit quantisation (q8_0).
gemma4:12b-it-bf1624GBThe 16-bit weights (bf16).
gemma4:31b-it-qat19GBA quantisation-aware training build of the 31B model.
gemma4:31b-it-q4_K_M20GBA 4-bit quantisation (q4_K_M).
gemma4:31b-it-q8_034GBAn 8-bit quantisation (q8_0).
gemma4:31b-it-bf1663GBThe 16-bit weights (bf16).

Each size also has MLX tags such as mxfp8 and nvfp4, listed with their own sizes. Source: ollama.com/library/gemma4/tags and ai.google.dev/gemma/docs/core.

How much memory does a model need?

Ollama's own documentation gives no formula for weights and context together, so this page states none of its own. What the sources do say is this. The tag size is the download. Google's Gemma 4 table gives approximate memory to load the weights, including 20 per cent overhead, and says it excludes the context cache: at 4-bit, 17.5 GB for the 31B model and 2.9 GB for E2B. A larger context window needs more on top.

Ollama's context length page says the default context depends on video memory: 4k below 24 GiB, 32k from 24 to 48 GiB, and 256k at 48 GiB or more. It says tasks such as web search, agents and coding tools should use at least 64000 tokens, and that more context needs more memory. The FAQ still says the default is 4096 tokens, so set the value you need explicitly. After loading, run ollama ps and read the PROCESSOR column to see whether the model sits fully on the GPU.

Ollama's FAQ also describes a quantised key and value cache. OLLAMA_KV_CACHE_TYPE can be q8_0, which it says uses about half the memory of f16, or q4_0, about a quarter. It needs Flash Attention, and the effect on quality depends on the model.

How do you choose a model in five steps?

  1. 1. Name the task

    Chat and writing, coding, image input, tool calling, reasoning, or embeddings for search. Match it to the tags in the task table. For code, also see best AI coding models.

  2. 2. Set the memory budget

    Take the video memory or unified memory you can give the model, then subtract room for context. Pick the largest tag whose size shown fits with room to spare, and use the Google table for the Gemma family.

  3. 3. Check the context you need

    Read the context column for the tag and set it explicitly in Ollama. Agents and coding tools need much more than the small default on many machines.

  4. 4. Read the licence

    Open the licence file on the tag page. See the caveat below, because licences differ even inside one vendor's range.

  5. 5. Test on your own prompts

    Run ten real tasks through two or three candidates and compare. Use how to evaluate an open-source LLM for the method, and best local LLM for the wider field beyond Ollama's library.

What is the licence caveat?

The Ollama software is MIT-licensed, but every model carries its own licence, and the licences in the task table are not the same. Most of the ten show Apache License 2.0. Llama 3.2 shows a community licence agreement with an acceptable use policy. DeepSeek R1 shows MIT but notes that distilled models derive from other base models. embeddinggemma shows the Gemma Terms of Use, while Gemma 4 shows Apache 2.0, so two models from one family can carry different terms.

Read the licence file on the tag page you pull, rather than the family page alone. Licences can change between versions, and fine-tunes and community uploads carry their own. This page summarises what the pages displayed on 2026-10-07. It is not legal advice, so have counsel read the text before you ship a model inside a product. A cloud tag is also a different service with its own terms.

Where Swfte fits

You do not need Swfte to pull and run an Ollama model, and Swfte does not choose a model for you or rank them. Where several people or services share a runtime, Swfte Connect gives you one OpenAI-compatible API with bring-your-own-key, per-workspace routing rules, budgets and usage caps, content-policy detectors for secrets and personal data, and an audit event stream. These are Built. It lets you change which model a workload uses as a routing decision instead of a setting inside each client. A hosted gateway needs a network path to your Ollama server.

Swfte Cortex is a desktop app that uses local models through Ollama and LM Studio. That is Built. See Connect for the gateway.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

What is the best Ollama model?

There is none that is best for every job. The library pages show different strengths: reasoning and agents, coding, image input, small sizes for modest hardware, and embeddings for search. Choose by task, then by the largest size that fits your memory with room for context, then confirm the licence. The task and memory tables above are the starting point.

Which Ollama model should I run with 8 GB of memory?

Ollama's quickstart recommends 8 GB of available video memory or unified memory for Gemma 4 E2B, whose download it puts at about 7.2 GB. Smaller tags in the library tables are llama3.2:3b at 2.0GB, granite4:3b at 2.1GB and deepseek-r1:8b at 5.2GB. Download size is not total memory, because context needs extra.

Which Ollama models can read images?

The library tags gemma4, qwen3.5 and mistral-small3.2 with vision, and their tag tables list Text, Image as the input. Ollama's CLI docs show a Gemma 4 image prompt. Check the tag table for the exact size you pull, because the tag is shown for the family rather than each size.

Which Ollama models are for embeddings?

embeddinggemma and nomic-embed-text carry the embedding tag, and both list a 2K context window. embeddinggemma shows the Gemma Terms of Use and nomic-embed-text shows Apache License 2.0. Embedding models return vectors for retrieval rather than chat answers. See how to build a RAG system for how they are used.

Are Ollama models free for commercial use?

The Ollama software is MIT-licensed, but each model has its own licence, and they differ. Several of the ten here show Apache License 2.0, while Llama 3.2 shows a community licence agreement with an acceptable use policy. Read the licence on the exact tag page before using a model commercially. This is not legal advice.

What does a cloud tag on an Ollama model mean?

A cloud tag, such as gemma4:31b-cloud, runs the model on Ollama's servers instead of your machine, so your prompts leave your device. Ollama says it does not use prompts and responses for training. If you need prompts to stay local, pull a normal tag and consider turning cloud features off.

Choose by task, then put one governed API in front

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.