API reference
Ollama API: endpoints, OpenAI compatibility, streaming and API keys
A reference to the Ollama HTTP API as documented on 2026-10-07, with copyable curl and Python examples and a plain answer on API keys.
Ollama serves an HTTP API at http://localhost:11434/api, with an OpenAI-compatible subset at http://localhost:11434/v1. The local server does not require an API key: Ollama's docs say the local API does not require authentication. Only direct requests to Ollama Cloud at https://ollama.com/api need a key, sent as a bearer token. Responses stream by default as newline-delimited JSON, unless you set stream to false.
Last verified 2026-10-07. Sources are listed at the end of the page.
Which endpoints does the Ollama API have?
Paths and methods come from the OpenAPI file at docs.ollama.com/openapi.yaml and the docs pages. The base URL for a local server is http://localhost:11434.
| Endpoint | What it does, per the docs |
|---|---|
| POST /api/generate | Generates a response for a prompt. Streams by default. |
| POST /api/chat | Generates the next chat message in a conversation. Read the answer from message.content. Accepts tools and a think field. |
| POST /api/embed | Creates vector embeddings for the input text. |
| GET /api/tags | Lists models. Against ollama.com it lists cloud models. |
| GET /api/ps | Lists the models currently loaded in memory. |
| POST /api/show | Shows model details, including the thinking controls a model supports. |
| POST /api/create | Creates a model. Needs a local Ollama server. |
| POST /api/copy | Copies a model to a new name. |
| POST /api/pull | Pulls a model and streams progress by default. |
| POST /api/push | Pushes a model and streams progress by default. |
| DELETE /api/delete | Deletes a model. Needs a local Ollama server. |
| GET /api/version | Returns the Ollama version. |
| HEAD and POST /api/blobs/{digest} | Checks whether a blob exists, or pushes one, when creating models. |
| POST /v1/systemone | Answers typed questions with a local System One model. Needs Ollama 0.35.0 or later and is not yet on Ollama Cloud. |
| GET /api/usage and GET /api/balance | Cloud only, at https://ollama.com with an API key: usage statistics and remaining credits. Each allows 10 requests per minute per user. |
Ollama says its API is not strictly versioned but is expected to stay stable and backwards compatible, with rare deprecations announced in release notes.
What does the OpenAI-compatible API support, and what does it not?
Ollama says it supports a subset of the OpenAI API. The columns repeat the checklists on docs.ollama.com/api/openai-compatibility.
| Route | Supported | Not supported |
|---|---|---|
| POST /v1/chat/completions | Chat completions, streaming, JSON mode, reproducible outputs, vision with base64 images, tools, reasoning and thinking control, stream_options with include_usage, and fields such as temperature, top_p, max_tokens, seed, stop and response_format. | logprobs, tool_choice, logit_bias, user, n, and image URLs (use base64). |
| POST /v1/completions | Completions, streaming, JSON mode, reproducible outputs, suffix and the common sampling fields. | logprobs, best_of, echo, logit_bias, user and n. The prompt must be a single string. |
| POST /v1/embeddings | model, input as a string or an array of strings, encoding format and dimensions. | Arrays of tokens, arrays of token arrays, and user. |
| /v1/models and /v1/models/{model} | Listing and fetching models. The created field is when the model was last modified, and owned_by is the Ollama username, defaulting to library. | Nothing else is listed. |
| POST /v1/responses | The stateless flavour (added in Ollama 0.13.3): streaming, tools, reasoning summaries, instructions and max_output_tokens. | previous_response_id, conversation, truncation, and stateful requests. |
The docs also list an Anthropic-compatible Messages route and a hosted /v1/messages, which takes the bearer key as well. See docs.ollama.com/api/anthropic-compatibility.
How do you call the Ollama API? Worked examples
These examples come from the Ollama repository README and docs pages. Start the server first (the app, or ollama serve) and pull the model you name.
1. Pull a model
From the command line, ollama pull gemma4. Through the API, curl http://localhost:11434/api/pull -d '{"model": "gemma4"}'. Add "stream": false to receive one reply instead of progress lines.
2. Chat with curl
curl http://localhost:11434/api/chat -d '{"model": "gemma4", "messages": [{"role": "user", "content": "Why is the sky blue?"}], "stream": false}'. Read the answer from message.content.
3. Stream the reply
Leave out "stream": false. Ollama returns one JSON object per line (content type application/x-ndjson), each with a piece of the answer and done set to false, and the last with done set to true. Timing and token counts arrive in that final object.
4. Use the Python library
pip install ollama, then: from ollama import chat; response = chat(model='gemma4', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}]); print(response.message.content).
5. Use the OpenAI Python client
pip install openai, then: from openai import OpenAI; client = OpenAI(base_url='http://localhost:11434/v1/', api_key='ollama'); chat_completion = client.chat.completions.create(messages=[{'role': 'user', 'content': 'Say this is a test'}], model='gpt-oss:20b'); print(chat_completion.choices[0].message.content). The docs mark the key as required but ignored. Replace the model with one you have pulled.
6. Create embeddings and list models
curl http://localhost:11434/api/embed -d '{"model": "embeddinggemma", "input": "Generate embeddings for this text"}' returns an embeddings array. curl http://localhost:11434/api/tags lists the models you have, and curl http://localhost:11434/api/ps shows what is loaded.
7. Call Ollama Cloud with a key
curl https://ollama.com/api/chat -H "Authorization: Bearer $OLLAMA_API_KEY" -H "Content-Type: application/json" -d '{"model": "gemma4:31b", "messages": [{"role": "user", "content": "Why is the sky blue?"}], "stream": false}'. This sends your prompt to Ollama's servers.
Does the Ollama API need an API key?
The local server does not. The authentication page says the local API at http://localhost:11434 does not require authentication, and the introduction says local requests do not need a key. OpenAI client libraries insist on a key value, and the docs say to pass any string, because Ollama ignores it.
Ollama Cloud does. Direct cloud inference at https://ollama.com/api and https://ollama.com/v1 requires an API key. Create one at ollama.com/settings/keys, set it as OLLAMA_API_KEY, and send it in the Authorization: Bearer header. For the hosted Anthropic-compatible endpoint, x-api-key alone is not supported. Keys do not expire, and you revoke one in the key settings. Keep keys out of browser code and source control, which the docs also advise.
There is a third path. After you run ollama signin, the local server authenticates cloud requests for you, so a local call to a model such as gpt-oss:120b-cloud works without a key in your client. The docs do not say how that interacts with exposure. This page infers, and has not tested, that a signed-in server which other people can reach could let them run cloud models on your account. Treat that as a risk to check, and do not expose a signed-in server without authentication in front. See is Ollama safe.
If you want a key on a local server, the server will not enforce one. Put a reverse proxy or gateway in front that checks keys, and keep Ollama bound to localhost behind it.
How do streaming, errors and usage fields behave?
- Streaming: the stream field defaults to true on generate, chat, pull and push. Set it to false for one application/json reply.
- Errors: status codes include 400 for bad requests, 404 when a model is missing, 429 for rate limits, 500 for server errors and 502 when a cloud model cannot be reached. The body is JSON with an error property.
- Mid-stream errors: the status code cannot change once the stream has started, so an error arrives as a final line with an error property. Check every line.
- Usage fields: responses carry total_duration, load_duration, prompt_eval_count, prompt_eval_cached_count, prompt_eval_duration, eval_count and eval_duration. Durations are in nanoseconds, and streamed calls put them in the final object.
- Keeping a model loaded: models stay in memory for five minutes by default. The keep_alive field overrides it per request, and OLLAMA_KEEP_ALIVE sets it for the server.
- Context: the OpenAI-compatible route has no field for context size. Create a model with a larger num_ctx in a Modelfile, or set OLLAMA_CONTEXT_LENGTH when serving.
Where Swfte fits
Swfte Connect can sit in front of an Ollama endpoint as an OpenAI-compatible API. In code, Connect resolves a workspace's own credential with an optional base URL and treats models outside its catalogue as bring-your-own-key, which is the mechanism for an Ollama /v1 route. On top of that it provides routing and fallback chains, budgets and usage caps, content-policy detectors for secrets and personal data with a redact action, and an audit event stream. These are Built. Ask us to confirm the setup for your endpoint, because this page documents no click path for it.
Two limits apply. The gateway needs a network path to your Ollama server and cannot call localhost on a laptop. And Ollama's /v1 route lists fields such as tool_choice as unsupported, so test any client that depends on them against your endpoint. Connect's free dashboard also works with Ollama directly, without routing prompts through the Swfte gateway.
You do not need a gateway to call Ollama from your own code. It earns its place when several callers share one server and you want caps, redaction and an audit trail. See Connect, providers and BYOK and how to set up an LLM gateway.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Ollama API introduction. Base URLs, local versus cloud, and versioning.
- Ollama API authentication. No local authentication, cloud API keys and ollama signin.
- Ollama API streaming. Newline-delimited JSON and stream false.
- Ollama API usage fields. Timing and token count fields.
- Ollama API errors. Status codes and mid-stream errors.
- Ollama OpenAI compatibility. Supported and unsupported routes and fields, and Python examples.
- Ollama OpenAPI specification. The list of paths and methods, and the stream defaults.
- Ollama cloud usage and balance. Cloud-only endpoints, keys and rate limits.
- Ollama cloud. Cloud API key setup and data handling.
- Ollama quickstart. Local and cloud request examples.
- Ollama FAQ. Default bind address, keep_alive and context settings.
- Ollama repository README. The curl and Python examples.
Frequently asked questions
Does Ollama need an API key?
Not for the local server. Ollama's docs say the local API at localhost:11434 does not require authentication, and OpenAI client examples pass a placeholder key that is ignored. Direct requests to Ollama Cloud at ollama.com do need an API key, created at ollama.com/settings/keys and sent as a bearer token.
What port does the Ollama API use?
Port 11434. The server binds 127.0.0.1 port 11434 by default, so the native API is at http://localhost:11434/api and the OpenAI-compatible API at http://localhost:11434/v1. Change the bind address with the OLLAMA_HOST environment variable. Anything wider than localhost needs authentication in front of it.
Is the Ollama API compatible with the OpenAI API?
Partly. Ollama supports a subset: chat completions, completions, embeddings, models and a stateless responses route. Fields such as tool_choice, logprobs, n and image URLs are listed as unsupported. Test the exact calls your client makes, and use the native /api routes when you need Ollama-specific features such as keep_alive.
How do I call the Ollama API from Python?
Install the library with pip install ollama and call chat(model='gemma4', messages=[...]), reading response.message.content. Or use the OpenAI client with base_url set to http://localhost:11434/v1/ and any string as the API key. Pull the model first with ollama pull so the server has it.
Does the Ollama API stream responses?
Yes, by default on generate, chat, pull and push. Responses arrive as newline-delimited JSON, one object per line, with done set to true on the last. Set stream to false for a single JSON reply. Handle error lines too, because an error mid-stream arrives as a final line with an error property.
How do I list the models on my Ollama server?
Send GET http://localhost:11434/api/tags to list the models you have downloaded, or use ollama ls on the command line. GET /api/ps shows which models are loaded in memory right now. Against ollama.com, the tags endpoint lists cloud models, and the same /v1/models route lists models for OpenAI clients.