Local runtimes

LM Studio vs Ollama: which local LLM runtime should you run?

A side-by-side of two local runtimes on documented facts: install, licence, formats, API, headless use and who should pick which.

Pick LM Studio if you want a desktop app for finding, loading and chatting with models. Pick Ollama if you want an open-source command line tool and background service that scripts and other software call. Both run models on your own hardware, both serve a local OpenAI-compatible API, and both listen on the local machine only until you change a setting. The differences that matter are the licence, the model catalogue, headless use and concurrency defaults. Everything below comes from each project's own documentation, read on 2026-10-07.

Last verified 2026-10-07. Sources are listed at the end of the page.

How do LM Studio and Ollama differ on documented facts?

Each cell comes from the vendor page listed under sources. Where a vendor page was silent, the cell says so.

TopicLM StudioOllama
Install surfaceDesktop app for macOS, Windows and Linux (the Linux build is an AppImage). A headless version, llmster, installs with a shell script and is driven by the lms command.Desktop app or installer on macOS and Windows, an install script or tarballs on Linux, and an official ollama/ollama Docker image. The ollama command starts the server or runs it as a service.
Licence and termsProprietary software under LM Studio's app terms. They grant a licence "solely for Your personal and / or internal business purposes" and bar sublicensing, redistribution, use as a service bureau or software-as-a-service, and reverse engineering. Some features are paid.MIT licence for the software, per the licence file in the repository. Model weights you pull carry their own licences.
Model formatsGGUF models through llama.cpp on Mac, Windows and Linux. MLX models as well on Apple silicon Macs.Imports GGUF and Safetensors weights through a Modelfile. The library also lists MLX tags, for example gemma4:31b-mlx, next to standard quantisation tags.
Catalogue sourceSearch and download from Hugging Face inside the app.The Ollama library at ollama.com/library, plus import of local files.
Local server and default portStart it from the Developer tab or with lms server start. Default port 1234, base URL http://localhost:1234/v1.Start it with ollama serve or the app. It binds 127.0.0.1 port 11434 by default; the OpenAI-compatible base URL is http://localhost:11434/v1.
OpenAI-compatible endpoints/v1/models, /v1/responses, /v1/chat/completions, /v1/embeddings and /v1/completions. An Anthropic-compatible Messages endpoint and a native REST API (listed as beta) are documented too./v1/chat/completions, /v1/completions, /v1/models, /v1/embeddings and a stateless /v1/responses. The docs list tool_choice, logprobs and image URLs as not supported. Native /api endpoints and an Anthropic compatibility page also exist.
AuthenticationNone by default. API tokens that every request must carry can be switched on (LM Studio 0.4.0 or newer).The local API at localhost:11434 does not require authentication. API keys apply only to direct access to Ollama Cloud.
Network exposureA Serve on Local Network option, or lms server start --bind 0.0.0.0. The docs recommend enabling authentication for any bind other than 127.0.0.1.The OLLAMA_HOST environment variable changes the bind address. The docs show examples for reverse proxies and tunnels that do not add authentication.
Concurrency defaultsMax Concurrent Predictions, default 4, uses continuous batching on the llama.cpp engine. The page says MLX support is coming soon.OLLAMA_NUM_PARALLEL defaults to 1 per the FAQ, and OLLAMA_MAX_QUEUE to 512. Several models can load at once if memory allows.
Hardware named in the docsmacOS 14 or newer on Apple silicon, 16 GB or more of RAM recommended. Windows x64 with AVX2 or ARM, 16 GB RAM and 4 GB VRAM recommended. Linux x64 or ARM64 on Ubuntu 20.04 or newer.macOS 14 or newer on Apple M series (CPU and GPU) or x86 (CPU only). NVIDIA GPUs with compute capability 5.0 or higher, AMD GPUs through ROCm v7, Vulkan, and Metal on Apple GPUs.
What leaves the machinePer its privacy policy, local messages, chat histories and documents are not transmitted. Update checks and model search or download do contact LM Studio. Cloud features are optional and paid.Per its privacy policy, no prompts, responses or other content processed locally are collected. Limited device and usage metadata (such as app version and request counts) and model download metadata may be. Cloud models send your prompts to Ollama.

What changes in client code when you switch runtimes?

Very little, if the client uses an OpenAI-compatible SDK. Change the base URL from http://localhost:1234/v1 (LM Studio) to http://localhost:11434/v1 (Ollama), or the reverse, and change the model name. Model names are not portable: the Ollama library uses tags such as gemma4:e4b, while LM Studio's docs use identifiers such as openai/gpt-oss-20b.

Two details catch people out. Ollama's docs say its client examples need an API key value but ignore it, so any placeholder string works against a local Ollama server. LM Studio needs a real token only if you have switched authentication on.

The other detail is feature coverage. Ollama lists tool_choice, logprobs, logit_bias, n and image URLs as unsupported on chat completions, and its responses endpoint is stateless. LM Studio documents its own list of compatible endpoints and a separate native REST API. If your client depends on one of those fields, test it on both before you commit.

Who should pick LM Studio, and who should pick Ollama?

Pick LM Studio when

A person wants to browse models, load them and chat in a window, with document chat and MCP servers built in. It also fits when you need MLX on a Mac, or an app you can run headless as llmster. Read the terms first if the use is anything other than internal work.

Pick Ollama when

Other software calls the model, you want an MIT-licensed tool, or you need a service you can run in Docker or under systemd. Its native /api and OpenAI-compatible /v1 surfaces are documented in one place. See the Ollama API guide.

Use both when

Different people have different habits. Because both expose OpenAI-compatible endpoints, a client written against one usually needs only a new base URL and model name for the other. Test the endpoint your client actually calls, because the supported fields differ.

Look elsewhere when

Many people or services need the model at once, or you serve large models across several GPUs. Both tools are built around a single machine. See alternatives to Ollama for engines built for shared serving.

What do the two licences allow for work use?

This is a plain reading of the published text, not legal advice. Have counsel read the current terms before you roll either tool out.

  • LM Studio: the grant is for "personal and / or internal business purposes", so using it inside your own organisation for work fits the wording. Hosting it for customers as a service, redistributing it or bundling it into your product does not. The terms are at lmstudio.ai/app-terms.
  • Ollama: the repository carries the MIT licence, which permits use, copying, modification and distribution on the standard MIT conditions. That covers the software only.
  • Models are separate: each model has its own licence. The library pages for gemma4 and gpt-oss show Apache 2.0, while llama3.2 shows the Llama 3.2 Community Licence Agreement. See best Ollama models.
  • Ollama Cloud is a separate service with its own plans and terms. A cloud model is not local. See Ollama Cloud pricing.

How do you choose in an afternoon?

  1. 1. Decide who calls the model

    One person at a desk points toward the app that suits them. A script, an editor plugin or another service points toward the one with the API you will call. Write down which endpoint the client uses.

  2. 2. Read the terms against your use

    Internal use is the common case for both. If the plan is to ship the runtime to customers or host it for them, check the LM Studio terms and the licences of the models.

  3. 3. Match the hardware lines

    Compare your machine with the hardware row above. Both document Apple silicon. Ollama's docs also list supported NVIDIA and AMD GPUs by family, while LM Studio's system requirements page gives memory guidance and names no GPU vendors.

  4. 4. Run your own prompts on both

    Load the same model on each and send the same ten real prompts through the API your client uses. Compare answers, memory use and start-up time on your machine rather than on a published chart.

  5. 5. Check concurrency and exposure

    If more than a few callers will hit it, read the concurrency row again. If anyone other than you needs network access, plan authentication first. See is Ollama safe.

Where Swfte fits

Swfte Connect is one OpenAI-compatible API with bring-your-own-key, routing and fallback chains, budgets and usage caps, content-policy detectors for secrets and personal data with a redact action, and an audit event stream. These are Built. Connect's free dashboard works with Ollama, vLLM, LM Studio or your own provider key, and its SDK sends telemetry without routing prompts through the Swfte gateway. See Connect.

Swfte Cortex is a desktop app with on-device knowledge bases, MCP servers and tool approvals, and it uses local models through both Ollama and LM Studio. That is Built.

You do not need Swfte to run either runtime. If one person uses one machine, neither tool needs a gateway. Connect becomes useful when several people or services share a runtime and you want budgets, redaction and an audit trail that neither runtime documents. A hosted gateway can reach only endpoints it can connect to over the network, not localhost on a laptop.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

Is LM Studio free for commercial use?

LM Studio's terms grant a licence solely for personal and internal business purposes, so using it inside your own organisation for work fits the wording. Selling it, hosting it for others as a service, or redistributing it is excluded, and some features are paid. Read the current terms at lmstudio.ai/app-terms. This is not legal advice.

Is Ollama free for commercial use?

The Ollama software is MIT-licensed according to the licence file in its repository, which permits commercial use on the standard MIT conditions. The models you pull carry separate licences, and Ollama Cloud is a separate paid service. Check the licence on each model page before you build a product on it.

Which one has the OpenAI-compatible server?

Both do. LM Studio serves on port 1234 and Ollama on port 11434, and both list chat completions, completions, embeddings, models and responses routes. The supported fields differ, so test the exact calls your client makes. Ollama's docs list tool_choice and logprobs as not supported on chat completions.

Can I run LM Studio or Ollama headless on a server?

Yes. LM Studio's headless version is called llmster and installs with a shell script, and its docs include a systemd unit. Ollama runs with ollama serve, as a systemd service on Linux, or in its official Docker image. For many simultaneous users, read the concurrency row and consider engines built for shared serving.

Is it safe to expose either one to my network?

Only with controls in front. Both listen locally by default. LM Studio can require API tokens, and its docs recommend authentication for any non-local bind. Ollama's docs say its local API needs no authentication, so put an authenticating proxy or gateway in front before anyone else can reach it.

Do they send my prompts to the vendor?

For local models, both vendors' privacy policies say prompt content stays on your machine. Both still contact their vendors for some things: LM Studio for update checks and model search or download, Ollama for limited usage metadata and model downloads. Cloud features on either side do send prompts to the vendor.

Put one governed API in front of whichever runtime you pick

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.