API reference

Ollama API: endpoints, OpenAI compatibility, streaming and API keys

A reference to the Ollama HTTP API as documented on 2026-10-07, with copyable curl and Python examples and a plain answer on API keys.

Ollama serves an HTTP API at http://localhost:11434/api, with an OpenAI-compatible subset at http://localhost:11434/v1. The local server does not require an API key: Ollama's docs say the local API does not require authentication. Only direct requests to Ollama Cloud at https://ollama.com/api need a key, sent as a bearer token. Responses stream by default as newline-delimited JSON, unless you set stream to false.

Last verified 2026-10-07. Sources are listed at the end of the page.

Which endpoints does the Ollama API have?

Paths and methods come from the OpenAPI file at docs.ollama.com/openapi.yaml and the docs pages. The base URL for a local server is http://localhost:11434.

EndpointWhat it does, per the docs
POST /api/generateGenerates a response for a prompt. Streams by default.
POST /api/chatGenerates the next chat message in a conversation. Read the answer from message.content. Accepts tools and a think field.
POST /api/embedCreates vector embeddings for the input text.
GET /api/tagsLists models. Against ollama.com it lists cloud models.
GET /api/psLists the models currently loaded in memory.
POST /api/showShows model details, including the thinking controls a model supports.
POST /api/createCreates a model. Needs a local Ollama server.
POST /api/copyCopies a model to a new name.
POST /api/pullPulls a model and streams progress by default.
POST /api/pushPushes a model and streams progress by default.
DELETE /api/deleteDeletes a model. Needs a local Ollama server.
GET /api/versionReturns the Ollama version.
HEAD and POST /api/blobs/{digest}Checks whether a blob exists, or pushes one, when creating models.
POST /v1/systemoneAnswers typed questions with a local System One model. Needs Ollama 0.35.0 or later and is not yet on Ollama Cloud.
GET /api/usage and GET /api/balanceCloud only, at https://ollama.com with an API key: usage statistics and remaining credits. Each allows 10 requests per minute per user.

Ollama says its API is not strictly versioned but is expected to stay stable and backwards compatible, with rare deprecations announced in release notes.

What does the OpenAI-compatible API support, and what does it not?

Ollama says it supports a subset of the OpenAI API. The columns repeat the checklists on docs.ollama.com/api/openai-compatibility.

RouteSupportedNot supported
POST /v1/chat/completionsChat completions, streaming, JSON mode, reproducible outputs, vision with base64 images, tools, reasoning and thinking control, stream_options with include_usage, and fields such as temperature, top_p, max_tokens, seed, stop and response_format.logprobs, tool_choice, logit_bias, user, n, and image URLs (use base64).
POST /v1/completionsCompletions, streaming, JSON mode, reproducible outputs, suffix and the common sampling fields.logprobs, best_of, echo, logit_bias, user and n. The prompt must be a single string.
POST /v1/embeddingsmodel, input as a string or an array of strings, encoding format and dimensions.Arrays of tokens, arrays of token arrays, and user.
/v1/models and /v1/models/{model}Listing and fetching models. The created field is when the model was last modified, and owned_by is the Ollama username, defaulting to library.Nothing else is listed.
POST /v1/responsesThe stateless flavour (added in Ollama 0.13.3): streaming, tools, reasoning summaries, instructions and max_output_tokens.previous_response_id, conversation, truncation, and stateful requests.

The docs also list an Anthropic-compatible Messages route and a hosted /v1/messages, which takes the bearer key as well. See docs.ollama.com/api/anthropic-compatibility.

How do you call the Ollama API? Worked examples

These examples come from the Ollama repository README and docs pages. Start the server first (the app, or ollama serve) and pull the model you name.

  1. 1. Pull a model

    From the command line, ollama pull gemma4. Through the API, curl http://localhost:11434/api/pull -d '{"model": "gemma4"}'. Add "stream": false to receive one reply instead of progress lines.

  2. 2. Chat with curl

    curl http://localhost:11434/api/chat -d '{"model": "gemma4", "messages": [{"role": "user", "content": "Why is the sky blue?"}], "stream": false}'. Read the answer from message.content.

  3. 3. Stream the reply

    Leave out "stream": false. Ollama returns one JSON object per line (content type application/x-ndjson), each with a piece of the answer and done set to false, and the last with done set to true. Timing and token counts arrive in that final object.

  4. 4. Use the Python library

    pip install ollama, then: from ollama import chat; response = chat(model='gemma4', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}]); print(response.message.content).

  5. 5. Use the OpenAI Python client

    pip install openai, then: from openai import OpenAI; client = OpenAI(base_url='http://localhost:11434/v1/', api_key='ollama'); chat_completion = client.chat.completions.create(messages=[{'role': 'user', 'content': 'Say this is a test'}], model='gpt-oss:20b'); print(chat_completion.choices[0].message.content). The docs mark the key as required but ignored. Replace the model with one you have pulled.

  6. 6. Create embeddings and list models

    curl http://localhost:11434/api/embed -d '{"model": "embeddinggemma", "input": "Generate embeddings for this text"}' returns an embeddings array. curl http://localhost:11434/api/tags lists the models you have, and curl http://localhost:11434/api/ps shows what is loaded.

  7. 7. Call Ollama Cloud with a key

    curl https://ollama.com/api/chat -H "Authorization: Bearer $OLLAMA_API_KEY" -H "Content-Type: application/json" -d '{"model": "gemma4:31b", "messages": [{"role": "user", "content": "Why is the sky blue?"}], "stream": false}'. This sends your prompt to Ollama's servers.

Does the Ollama API need an API key?

The local server does not. The authentication page says the local API at http://localhost:11434 does not require authentication, and the introduction says local requests do not need a key. OpenAI client libraries insist on a key value, and the docs say to pass any string, because Ollama ignores it.

Ollama Cloud does. Direct cloud inference at https://ollama.com/api and https://ollama.com/v1 requires an API key. Create one at ollama.com/settings/keys, set it as OLLAMA_API_KEY, and send it in the Authorization: Bearer header. For the hosted Anthropic-compatible endpoint, x-api-key alone is not supported. Keys do not expire, and you revoke one in the key settings. Keep keys out of browser code and source control, which the docs also advise.

There is a third path. After you run ollama signin, the local server authenticates cloud requests for you, so a local call to a model such as gpt-oss:120b-cloud works without a key in your client. The docs do not say how that interacts with exposure. This page infers, and has not tested, that a signed-in server which other people can reach could let them run cloud models on your account. Treat that as a risk to check, and do not expose a signed-in server without authentication in front. See is Ollama safe.

If you want a key on a local server, the server will not enforce one. Put a reverse proxy or gateway in front that checks keys, and keep Ollama bound to localhost behind it.

How do streaming, errors and usage fields behave?

  • Streaming: the stream field defaults to true on generate, chat, pull and push. Set it to false for one application/json reply.
  • Errors: status codes include 400 for bad requests, 404 when a model is missing, 429 for rate limits, 500 for server errors and 502 when a cloud model cannot be reached. The body is JSON with an error property.
  • Mid-stream errors: the status code cannot change once the stream has started, so an error arrives as a final line with an error property. Check every line.
  • Usage fields: responses carry total_duration, load_duration, prompt_eval_count, prompt_eval_cached_count, prompt_eval_duration, eval_count and eval_duration. Durations are in nanoseconds, and streamed calls put them in the final object.
  • Keeping a model loaded: models stay in memory for five minutes by default. The keep_alive field overrides it per request, and OLLAMA_KEEP_ALIVE sets it for the server.
  • Context: the OpenAI-compatible route has no field for context size. Create a model with a larger num_ctx in a Modelfile, or set OLLAMA_CONTEXT_LENGTH when serving.

Where Swfte fits

Swfte Connect can sit in front of an Ollama endpoint as an OpenAI-compatible API. In code, Connect resolves a workspace's own credential with an optional base URL and treats models outside its catalogue as bring-your-own-key, which is the mechanism for an Ollama /v1 route. On top of that it provides routing and fallback chains, budgets and usage caps, content-policy detectors for secrets and personal data with a redact action, and an audit event stream. These are Built. Ask us to confirm the setup for your endpoint, because this page documents no click path for it.

Two limits apply. The gateway needs a network path to your Ollama server and cannot call localhost on a laptop. And Ollama's /v1 route lists fields such as tool_choice as unsupported, so test any client that depends on them against your endpoint. Connect's free dashboard also works with Ollama directly, without routing prompts through the Swfte gateway.

You do not need a gateway to call Ollama from your own code. It earns its place when several callers share one server and you want caps, redaction and an audit trail. See Connect, providers and BYOK and how to set up an LLM gateway.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

Does Ollama need an API key?

Not for the local server. Ollama's docs say the local API at localhost:11434 does not require authentication, and OpenAI client examples pass a placeholder key that is ignored. Direct requests to Ollama Cloud at ollama.com do need an API key, created at ollama.com/settings/keys and sent as a bearer token.

What port does the Ollama API use?

Port 11434. The server binds 127.0.0.1 port 11434 by default, so the native API is at http://localhost:11434/api and the OpenAI-compatible API at http://localhost:11434/v1. Change the bind address with the OLLAMA_HOST environment variable. Anything wider than localhost needs authentication in front of it.

Is the Ollama API compatible with the OpenAI API?

Partly. Ollama supports a subset: chat completions, completions, embeddings, models and a stateless responses route. Fields such as tool_choice, logprobs, n and image URLs are listed as unsupported. Test the exact calls your client makes, and use the native /api routes when you need Ollama-specific features such as keep_alive.

How do I call the Ollama API from Python?

Install the library with pip install ollama and call chat(model='gemma4', messages=[...]), reading response.message.content. Or use the OpenAI client with base_url set to http://localhost:11434/v1/ and any string as the API key. Pull the model first with ollama pull so the server has it.

Does the Ollama API stream responses?

Yes, by default on generate, chat, pull and push. Responses arrive as newline-delimited JSON, one object per line, with done set to true on the last. Set stream to false for a single JSON reply. Handle error lines too, because an error mid-stream arrives as a final line with an error property.

How do I list the models on my Ollama server?

Send GET http://localhost:11434/api/tags to list the models you have downloaded, or use ollama ls on the command line. GET /api/ps shows which models are loaded in memory right now. Against ollama.com, the tags endpoint lists cloud models, and the same /v1/models route lists models for OpenAI clients.

Add caps, redaction and audit in front of your Ollama endpoint

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.