Pricing and sovereignty
Ollama Cloud pricing: plans, usage credits and what it sends
The plans, prices and usage rules Ollama published on 2026-10-07, the cloud models and rates listed, and the sovereignty point that cloud prompts leave your machine.
Ollama's pricing page, read on 2026-10-07, lists Free at $0, Pro at $20 a month ($200 a year billed annually), Max at $100 a month, Team at $500 a month, and Enterprise as custom. Usage is measured in tokens at each model's rate and drawn from dollar credits. A cloud model sends your prompts to Ollama's servers, so it is no longer local. Ollama says it does not train on prompts and responses. Prices and limits can change, so confirm them on ollama.com/pricing.
Last verified 2026-10-07. Sources are listed at the end of the page.
What is Ollama Cloud?
Ollama's docs describe it as running models in Ollama's cloud from your apps or terminal, with no model or app download required. You call it in one of two ways. Directly, with an API key, at https://ollama.com/api or the OpenAI-compatible https://ollama.com/v1. Or through the Ollama app or CLI after you sign in, using a tag with the cloud suffix, for example gemma4:cloud. The API names cloud models by the names that https://ollama.com/api/tags returns, such as gemma4:31b.
Cloud models use their maximum context length by default. Ollama's usage settings show upcoming model retirements, and the docs say downloaded local models are not affected by them. For the API details, see the Ollama API guide.
What do Ollama Cloud plans cost?
Prices are as published on ollama.com/pricing on 2026-10-07, in dollars as the page shows them. The page does not state tax treatment in what was read.
| Plan | Price as published | What the page lists | Concurrent requests |
|---|---|---|---|
| Free | $0 | Run models locally, starter usage credits included, access to starter models, credits can be added to use all models, no service fees. | 1 |
| Pro | $20 a month, or $200 a year billed annually ($16.67 a month billed annually) | $60 of usage credits per month, access to larger pro models, run multiple models concurrently, fast mode (shown as coming soon). | 3 |
| Max | $100 a month | Everything in Pro, plus $300 of usage credits per month, early access to the newest models and 10 concurrent requests. | 10 |
| Team | $500 a month (labelled early access) | Unlimited users, $1,000 of usage credits per month shared across the team, centralised billing and administration, priority support, shared projects, skills and instructions (shown as coming soon). | 10 |
| Enterprise | Custom, with volume usage pricing | Everything in Team, plus model access controls, cost budgets for users and API keys, a private Slack channel with dedicated support, and custom security questionnaires. | Not stated on the pricing page |
Source: ollama.com/pricing, read on 2026-10-07. Check the page before you buy, because plans, credits and limits may change.
Which models run in Ollama Cloud, and at what rates?
The pricing page lists these models, with prices in dollars per million tokens. Its notes say off-peak pricing applies outside 12:00 and 18:00 UTC on weekdays and all day on weekends, and it shows off-peak rows only for the two DeepSeek models. A dash means the page shows no cached-input price.
| Model | Input | Cached input | Output |
|---|---|---|---|
| deepseek-v4.1-flash | $0.30 | $0.006 | $1.20 |
| deepseek-v4.1-flash (Off-Peak) | $0.15 | $0.003 | $0.60 |
| deepseek-v4-pro | $1.32 | $0.044 | $3.96 |
| deepseek-v4-pro (Off-Peak) | $0.66 | $0.022 | $1.98 |
| gemma4 | $0.14 | $0.05 | $0.40 |
| glm-5.3 | $1.40 | $0.26 | $4.40 |
| glm-5.3-flash | $0.15 | $0.03 | $0.50 |
| glm-5.2 | $1.40 | $0.26 | $4.40 |
| gpt-oss:120b | $0.15 | $0.014 | $0.60 |
| gpt-oss:20b | $0.07 | $0.035 | $0.30 |
| kimi-k3 | $3.00 | $0.30 | $15.00 |
| kimi-k2.7-code | $0.95 | $0.19 | $4.00 |
| kimi-k2.6 | $0.95 | $0.16 | $4.00 |
| minimax-m3 | $0.60 | $0.12 | $2.40 |
| minimax-m2.7 | $0.30 | $0.06 | $1.20 |
| mistral-large-4 | $0.68 | $0.07 | $2.09 |
| mistral-large-3 | $0.50 | - | $1.50 |
| nemotron-3-nano | $0.06 | - | $0.24 |
| nemotron-3-super | $0.015 | $0.015 | $0.60 |
| nemotron-3-ultra | $0.10 | $0.10 | $3.00 |
Source: the model pricing table on ollama.com/pricing, read on 2026-10-07. Which models are free-plan starter models is not named on the page.
How does Ollama describe its usage limits?
- Credits, not request counts. Pro, Max and Team include a set dollar amount of usage credits each month. Usage is measured in tokens at each model's rates. Free includes a starter amount for a smaller set of starter models, and credits can be bought on any plan.
- Order of use. Included plan credits are used first, then the purchased credit balance. Team usage draws from one shared balance.
- Reset and rollover. Included usage resets monthly on the day the subscription started, including annual plans, and monthly from sign-up on Free. Unused included usage does not roll over.
- Warnings. On paid plans Ollama emails a reminder at 90 per cent of included monthly usage, and you can turn that off. The usage and balance endpoints in the API let you poll your own figures, at 10 requests per minute per user.
- Concurrency. Free allows 1 concurrent request, Pro 3, and Max and Team 10. Requests beyond the limit are queued, up to a fixed limit, and rejected once the queue is full.
- Older plans. The page says older Pro and Max plans had session and weekly limits, that switching resets usage to the new plan's full monthly amount, and that the old limits then stop applying.
- Your own hardware. The page says running models on your own hardware is always unlimited. It means no usage meter, not no cost.
Is a cloud model still local? What happens to your prompts?
No. A cloud model runs on Ollama's servers, so the prompt and the answer travel there. It is not on-device and it is not on your infrastructure. Ollama's pricing page says it hosts models and compute primarily in the United States, and that to serve global demand it may route to Europe and Singapore for additional capacity. Its privacy policy says data may be transferred to and processed in the United States.
On retention, Ollama's statements read as follows. The privacy policy says it processes cloud prompts and responses transiently, that they are not stored beyond the time needed to fulfil the request, that technical measures are designed to minimise retention, and that it does not train on them. The FAQ says it does not store or log that content and never trains on it. The pricing page says prompt or response data is never logged or trained on, that it hosts with NVIDIA Cloud Providers, and that it requires no logging, no training and zero data retention from partners.
These are the vendor's own statements. The wording differs between pages, with "never logged" in one place and "transient processing" in another. Swfte has not audited them. Ollama also collects account information and limited usage metadata. If your data has residency rules, sector rules or a client contract that names where processing may happen, a US-primary service that may route elsewhere changes the analysis. That is a fact about location, not legal advice.
Turning it off is possible. Set OLLAMA_NO_CLOUD=1, or disable_ollama_cloud in ~/.ollama/server.json, and restart. You lose cloud models and web search. See is Ollama safe.
How do you decide between Ollama Cloud and your own hardware?
1. Classify the prompts
Sort what you will send into public, internal and restricted. Restricted data, and anything under a residency or confidentiality commitment, should stay on infrastructure you control.
2. Check whether the same model runs locally
Two models in the price table, gemma4 and gpt-oss, are also in the Ollama library with Apache License 2.0 text, so the same weights can run on your own hardware. The library lists gpt-oss:20b at 14GB and gpt-oss:120b at 65GB. See run Gemma 4 on Ollama.
3. Compare cost on your own numbers
Cloud is priced per million tokens, as above. Local is priced by hardware, power and your time. Ollama publishes no hardware price or break-even figure, and neither does this page. Get quotes and see cloud GPU options if you want to rent instead.
4. Decide per workload, not per company
Burst or non-sensitive work may suit the cloud. Steady, sensitive work usually belongs on your own hardware or a dedicated environment.
5. Enforce the decision
A rule that lives in a policy document does not stop a developer from signing in. Route model calls through one place where the rule is applied, and switch cloud features off on shared servers.
Where Swfte fits
Swfte is not a GPU cloud and does not resell Ollama Cloud. Sovereignty comes from choosing where each prompt may go, and bring-your-own-key keeps that choice with you. Swfte Connect is one OpenAI-compatible API with bring-your-own-key and routing, fallback chains, budgets and usage caps, content-policy detectors for secrets and personal data with a redact action, and an audit event stream. These are Built. With Connect in the path, a workload that handles restricted data can be routed to your own runtime, and a policy can redact secrets and personal data before a prompt reaches any outside endpoint.
The gateway only controls calls that go through it. It cannot stop a developer from signing a laptop in to Ollama Cloud directly. Pair it with shadow AI discovery. See Connect and providers and BYOK.
If every prompt you send is non-sensitive and the published price suits you, you do not need any of this.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Ollama pricing. Plans, prices, included credits, concurrency, model rates, hosting regions and data statements.
- Ollama privacy policy. Cloud processing, retention, metadata and United States transfer statements.
- Ollama cloud docs. What Ollama Cloud is, API key setup, model names and data handling.
- Ollama FAQ. The prompt data statement and the switch to disable cloud features.
- Ollama API authentication. Cloud API keys and signing in.
- Ollama cloud balance endpoint. Included and purchased credits, and the legacy session and weekly limits.
- Ollama cloud usage endpoint. Usage statistics and rate limits.
- Ollama library: gpt-oss. Local sizes and Apache 2.0 for a model also offered in the cloud.
- Ollama library: gemma4. Local sizes, the cloud tags and Apache licence text.
Frequently asked questions
How much does Ollama Cloud cost?
On 2026-10-07 the pricing page listed Free at $0, Pro at $20 a month or $200 a year billed annually, Max at $100 a month, Team at $500 a month, and Enterprise as custom. Pro includes $60 of usage credits a month, Max $300 and Team $1,000 shared. Credits can also be bought. Check ollama.com/pricing for current figures.
Is Ollama Cloud free?
There is a Free plan at $0 with starter usage credits and access to starter models, and the pricing page says running models locally is always unlimited. Using all cloud models needs credits, which any plan can buy. The page does not name which models are the starter models, so check your account.
Which models are available on Ollama Cloud?
The pricing page lists 18 models, including gemma4, gpt-oss:120b, gpt-oss:20b, several DeepSeek, GLM, Kimi, MiniMax and Mistral models, and three Nemotron models, each with input, cached input and output prices per million tokens. The API lists the current set at https://ollama.com/api/tags, and the list can change.
Does Ollama store or train on my cloud prompts?
Ollama says no to training. Its privacy policy says cloud prompts and responses are processed transiently and not stored beyond the time needed to fulfil the request, and its pricing page says prompt or response data is never logged or trained on. These are the vendor's statements, with different wording on different pages, and Swfte has not audited them.
Where does Ollama Cloud process my data?
Ollama's pricing page says it hosts models and compute primarily in the United States and may route to Europe and Singapore for additional capacity. The privacy policy says data may be transferred to and processed in the United States. If you have residency requirements, read both pages and ask Ollama directly before sending data.
How do I use Ollama without the cloud?
Run local models and turn cloud features off. Set OLLAMA_NO_CLOUD=1, or set disable_ollama_cloud to true in ~/.ollama/server.json, then restart Ollama. You lose cloud models and web search, and the logs then show that cloud is disabled. Local models stay free of any usage meter, because they run on your hardware.