Self-host how-to

Run Gemma 4 on Ollama: install, pull, run and call the API

A how-to for running Gemma 4 locally with Ollama, with sizes, licence, context window and memory taken from Ollama and Google pages.

Gemma 4 is the current Gemma generation on Google's Gemma page and in the Ollama library on 2026-10-07. Install Ollama, run ollama pull gemma4 (or a size tag such as gemma4:e4b), then ollama run gemma4 or call the local API on port 11434. Google's model card gives Apache 2.0 as the licence and lists five sizes: E2B, E4B, 12B, 26B A4B and 31B. The small models have a 128K context window and the larger ones 256K. Pick the size by memory, because the downloads are larger than the parameter names suggest.

Last verified 2026-10-07. Sources are listed at the end of the page.

Which Gemma 4 size should you pull?

Ollama sizes are the download sizes in the library tag table, where a range spans the formats listed for that tag. Google sizes are its approximate memory to load the weights, including 20 per cent overhead, and exclude the context cache.

Ollama tagDownload size shown by OllamaContext windowGoogle estimate at 4-bit (Q4_0)Google estimate at 16-bit (BF16)
gemma4:e2b4.6GB to 7.5GB128K2.9 GB11.4 GB
gemma4:e4b (the latest tag)6.6GB to 9.5GB128K4.5 GB17.9 GB
gemma4:12b7.7GB to 8.0GB256K6.7 GB26.7 GB
gemma4:26b (26B A4B)16GB to 19GB256K14.4 GB57.7 GB
gemma4:31b19GB to 20GB256K17.5 GB69.9 GB

Sources: ollama.com/library/gemma4 and ai.google.dev/gemma/docs/core, read on 2026-10-07. Google says its numbers may change with the inference tool.

How do you install Ollama, pull Gemma 4 and run it?

  1. 1. Install Ollama

    Download the app from ollama.com/download, or use the commands in the repository README: curl -fsSL https://ollama.com/install.sh | sh on macOS and Linux, and irm https://ollama.com/install.ps1 | iex on Windows. On Linux you can run ollama serve by hand or set up the systemd service described in the Linux docs. Read an install script before you pipe it to a shell.

  2. 2. Check your memory first

    Ollama's quickstart uses Gemma 4 E2B as its local example. It puts the download at about 7.2 GB and recommends 8 GB of available video memory or unified memory on a Mac, and warns that larger context windows need more. The library lists 4.6GB to 7.5GB for the same tag. Plan for the larger figure.

  3. 3. Pull the model

    Run ollama pull gemma4 for the default tag, which the library lists as e4b. Or choose a size, for example ollama pull gemma4:e2b, ollama pull gemma4:12b, ollama pull gemma4:26b or ollama pull gemma4:31b.

  4. 4. Run it

    Run ollama run gemma4 and type a prompt. For an image, Ollama's vision docs give ollama run gemma4 ./image.png what is in this image? as the quick start. List what you have with ollama ls.

  5. 5. Check where it loaded

    Run ollama ps. The PROCESSOR column shows 100% GPU when the model is fully on the GPU, or a CPU and GPU split when it spilled into system memory. The context-length docs show an example line for gemma4:latest with a 131072 context.

  6. 6. Call the API

    The local server answers on port 11434. The repository README gives: curl http://localhost:11434/api/chat -d '{"model": "gemma4", "messages": [{"role": "user", "content": "Why is the sky blue?"}], "stream": false}'. Leave out "stream": false to receive streamed lines. In Python, pip install ollama, then from ollama import chat, response = chat(model='gemma4', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}]), print(response.message.content). The Ollama API guide covers the rest.

What do Google and Ollama say about licence, context and modalities?

Which generation is current: Google's Gemma page lists Gemma 4 as the latest family, announced in April 2026 on that page, with follow-ups listed for May and June 2026 (multi-token prediction drafters, a 12B model and QAT builds). The Ollama library lists it as gemma4, shown as updated one week earlier when read.

Licence: Google's Gemma 4 model card gives Apache 2.0, and the Ollama tag pages show Apache License 2.0 text. Other Gemma-family models can differ. The embeddinggemma page in the Ollama library shows the Gemma Terms of Use, so check each tag. This is not legal advice.

Context and input: the model card gives 128K tokens for E2B and E4B and 256K for 12B, 26B A4B and 31B. All sizes take text and image input and produce text. Google says audio input is featured on E2B, E4B and 12B. The Ollama page shows an audio tag, but its tag table lists the input as Text, Image for every size, so this page does not state that audio input works through Ollama. Verify it before you rely on it.

Capabilities: Google describes configurable thinking modes, native function calling and system-role support. The Ollama page tags the family vision, tools and thinking. The 26B A4B model is a mixture of experts with 25.2B total and 3.8B active parameters.

What are the common pitfalls with Gemma 4 on Ollama?

  • The context window is smaller than the model's. Ollama's context-length page sets the default by video memory: 4k below 24 GiB, 32k from 24 to 48 GiB, 256k at 48 GiB or more. The FAQ still says 4096. Set OLLAMA_CONTEXT_LENGTH or num_ctx, and use at least 64000 tokens for agents and coding tools, per Ollama.
  • The download is bigger than the name suggests. In E2B and E4B the E means effective parameters, and the files also hold embeddings and vision or audio parts. Use the download size, not the number in the name.
  • The 26B model is not a 4B model in memory. Google says all 26 billion parameters must be loaded, so its memory sits close to a dense 26B model even though 3.8B are active per token.
  • Silent CPU offload. If ollama ps shows a CPU and GPU split, generation slows. Choose a smaller tag, a shorter context or a quantised cache.
  • Sampling defaults. Google recommends temperature 1.0, top_p 0.95 and top_k 64, and the Ollama tag ships those values as its parameters. Changing them without testing can change answers.
  • Thinking in multi-turn chats. Google and Ollama both say to keep earlier thought blocks out of the history and send only the final answers. Google makes an exception for tool-call turns, where thinking content should be kept. For images, place the image before the text.
  • Parallel requests multiply memory. Ollama's FAQ says required memory scales with OLLAMA_NUM_PARALLEL times the context length, and idle models unload after five minutes by default.
  • Cloud tags are not local. gemma4:cloud and gemma4:31b-cloud run on Ollama's servers. See Ollama Cloud pricing.

How do you expose the server safely?

Ollama binds 127.0.0.1 port 11434 by default, and its docs say the local API needs no authentication. Treat any wider exposure as a decision.

  1. 1. Keep the default if only you use it

    Leave OLLAMA_HOST unset. Then only processes on the machine can reach the API.

  2. 2. Bind to a private interface if others must reach it

    Set OLLAMA_HOST, for example 0.0.0.0:11434 as in the FAQ, only on a private network or behind a firewall that limits who can connect. On a Mac app use launchctl setenv, on Linux systemctl edit ollama.service, and on Windows the user environment variables.

  3. 3. Put authentication in front

    The FAQ's nginx, ngrok and Cloudflare Tunnel examples forward requests and add no authentication, so a tunnel makes the unauthenticated API reachable from outside. Add a proxy or gateway that checks a key or identity before requests reach Ollama.

  4. 4. Mind browser access

    Ollama allows cross-origin requests from 127.0.0.1 and 0.0.0.0 by default. Add origins with OLLAMA_ORIGINS only for sites you trust.

  5. 5. Log and limit

    Record who calls what, and cap usage per caller. Ollama's docs describe no per-user limits for the local API. See is Ollama safe.

Where Swfte fits

You do not need Swfte to run Gemma 4 on your own machine. When several people or services share the server, Swfte Connect puts one OpenAI-compatible API in front, with bring-your-own-key and an optional base URL, budgets and usage caps, content-policy detectors for secrets and personal data with a redact action, and an audit event stream. These are Built. The gateway needs a network path to your Ollama endpoint, and Ollama's /v1 route does not support every OpenAI field. See Connect and providers and BYOK.

Swfte Cortex is a desktop app that can use local models through Ollama (Built). For the wider choice of models, see best Ollama models and best local LLM.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

Is Gemma 4 available on Ollama?

Yes. The Ollama library lists gemma4 with the sizes e2b, e4b, 12b, 26b and 31b, plus cloud tags. The default tag, latest, is e4b. Pull it with ollama pull gemma4 and run it with ollama run gemma4. The library tags it vision, tools and thinking, and its tag table lists Text, Image input.

How much memory does Gemma 4 need on Ollama?

It depends on the size. Ollama's library lists downloads from 4.6GB for E2B to 20GB for the 31B model, and its quickstart recommends 8 GB of available video memory for E2B. Google estimates weights alone at 2.9 GB for E2B up to 17.5 GB for 31B at 4-bit. Context needs more on top.

Is Gemma 4 free for commercial use?

Google's Gemma 4 model card lists the licence as Apache 2.0, and the Ollama tag pages show Apache License 2.0 text, which permits commercial use on its conditions, such as keeping the notices. Other Gemma models can have different terms, so check the tag you pull. This is not legal advice.

What is the context window of Gemma 4 on Ollama?

The models support 128K tokens for E2B and E4B and 256K for 12B, 26B A4B and 31B, per Google. Ollama's default context is lower and depends on your video memory, so set it yourself. Ollama advises at least 64000 tokens for agents, web search and coding tools, and larger context uses more memory.

Does Gemma 4 on Ollama support images and tool calling?

Image input and tools are tagged on the Ollama library page, and Google describes native function calling. Ollama's vision docs show Gemma 4 taking an image from the command line and through the API. Audio input is listed by Google for E2B, E4B and 12B, but this page did not verify it through Ollama.

Can I run Gemma 4 on Ollama without sending data to the cloud?

Yes, with the normal tags. Ollama's privacy policy says it does not collect prompts or responses processed locally. The cloud tags are different and send prompts to Ollama's servers. You can disable cloud features with OLLAMA_NO_CLOUD=1 or the disable_ollama_cloud setting in the server file.

Run Gemma 4 locally and keep control of who can reach it

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.