Serving engines

SGLang vs vLLM: which inference engine to run, by workload

Compare SGLang and vLLM on what each project documents, and pick one by workload instead of by a benchmark chart.

Both are open-source, Apache 2.0, OpenAI-compatible inference servers for GPUs you control, and both document prefix caching, speculative decoding, structured output, quantisation and multi-GPU parallelism. The documented differences are narrow: SGLang is built around RadixAttention and agentic workloads, vLLM around PagedAttention and the widest hardware and plugin list. This page publishes no throughput numbers, because the only way to know which is quicker for you is to replay your own traffic through both.

Last verified 2026-10-07. Sources are listed at the end of the page.

SGLang vs vLLM: what each project documents

Every cell below is taken from the project’s own documentation or repository on 2026-10-07. A blank in the project’s docs is written as “not stated”, never filled from memory.

TopicvLLMSGLang
LicenceApache-2.0 (repository).Apache 2.0 (repository). The docs say the project is hosted under the non-profit open-source organisation LMSYS.
Core techniques named by the projectPagedAttention, continuous batching with chunked prefill and prefix caching, optimised attention kernels (FlashAttention, FlashInfer and others).RadixAttention, prefix caching and multi-GPU parallelism (documentation home page). The SGLang paper describes RadixAttention as KV cache reuse.
Speculative decodingn-gram, suffix, EAGLE and DFlash variants listed.EAGLE-2 and EAGLE-3, MTP, UNO, DFLASH, a standalone draft model and an n-gram variant listed.
Structured outputxgrammar or guidance as backends; JSON schema, regex, grammar and choice; on by default in the OpenAI-compatible server.XGrammar (default), Outlines or Llguidance; JSON schema, regex or EBNF, one constraint per request; usable through response_format.
ParallelismTensor, pipeline, data, expert and context parallelism. Multi-node through Ray or multiprocessing.Tensor, pipeline, attention-context and MoE data parallelism, plus a multi-node flag for tensor parallelism. Data parallelism is launched through a router.
QuantisationFP8, MXFP8 and MXFP4, NVFP4, INT8, INT4, GPTQ and AWQ, GGUF, compressed-tensors and more.The server-arguments page lists awq, fp8, gptq, marlin, bitsandbytes and gguf, among others.
OpenAI-compatible APIChat completions, completions, responses, embeddings, transcriptions and translations. Anthropic Messages API and gRPC are also listed.Completions, vision and embeddings pages in the OpenAI-compatible section; the docs home page says it is compatible with Hugging Face and OpenAI APIs.
Hardware namedNVIDIA, AMD and Intel GPUs, x86, ARM and PowerPC CPUs, plus plugins for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend, Apple Silicon and others.NVIDIA, AMD Instinct, Google TPU, Intel GPUs and CPUs, Apple Silicon, Huawei Ascend and Moore Threads, with specific cards listed in the README.
Disaggregated prefill and decodeDocumented and marked experimental and subject to change. The page says it does not improve throughput.Documented with Mooncake and NIXL transfer engines and single-node and multi-node examples. The page lists unsupported features.
Metrics and health/metrics in Prometheus format and a /health endpoint, both listed./metrics when started with --enable-metrics, with time to first token, queue and token metrics. A health endpoint was not found on the pages read for this guide.

Sources are listed at the foot of the page. Both projects release often, so re-check the version you will install.

What actually differs between them?

The two engines overlap more than their marketing suggests. vLLM’s documentation names automatic prefix caching, and SGLang’s names RadixAttention. Both reuse KV cache across requests that share a prefix. vLLM’s design page says it hashes each KV-cache block by the tokens in the block and the tokens before it. SGLang’s paper describes RadixAttention as KV cache reuse. Neither page tells you which wins on your prompts.

Where the projects present themselves differently is emphasis. vLLM describes itself as a library for LLM inference and serving and leads with memory management and breadth: hardware plugins, many quantisation formats, and an API server that lists more than chat completions. SGLang describes itself as an inference framework optimised for agentic workloads, reinforcement-learning rollouts and large-scale serving,.

The two papers each report large speedups, but against the baselines of their day, and both engines have moved on. This page does not quote them as a comparison. If a vendor chart ranks one engine, ask for the model, the GPU, the prompt mix, the concurrency, the engine version and the date.

Which should I pick for my workload?

These are decision rules drawn from the documented features. They tell you which engine to test first. They do not predict a winner.

WorkloadTest firstWhy, and what to check
Shared multi-user chat on a common modelvLLM, then SGLangBoth serve the OpenAI API with continuous batching. Check that the exact model and quantisation you want is listed as supported by the version you install, and measure time to first token at your real concurrency.
Agent loops with long shared prefixesSGLang, then vLLMSGLang’s own pages lead with RadixAttention and agentic workloads. vLLM also has automatic prefix caching. Replay a real multi-turn trace through both and compare cache hit behaviour in their metrics.
Structured JSON output at volumeEitherBoth document xgrammar. Differences are in the secondary backends (Outlines and Llguidance for SGLang, guidance for vLLM) and in schema features. Test your own JSON schema, including its hardest case, on both.
Hardware you already ownThe one whose install page lists itBoth lists are long and differ at the edges: vLLM names Gaudi and IBM Spyre plugins, SGLang names Moore Threads. Check your exact accelerator and driver on each project’s install page before anything else.
Very large mixture-of-experts models across nodesWhichever documents your modelBoth document expert and tensor parallelism and multi-node launch. Model-specific recipes decide this, so read the model card’s serving section first.
Separating prefill and decodeSGLang, with careSGLang documents it with two transfer engines. vLLM marks its version experimental. Both pages list limits, and neither is a reason to adopt the idea before you have a latency target it would fix.

How do I compare them on my own traffic?

One afternoon of measurement is worth more than any published chart. Keep the setup identical for both engines.

  1. 1. Freeze the model artefact

    Use the same weights and quantisation for both. If one engine needs a different format, record that and treat the pair as two separate candidates.

  2. 2. Capture real prompts

    Export a few hundred anonymised requests with their real system prompts, tool definitions and conversation history. Prefix reuse only shows up when the prompts are real.

  3. 3. Replay at your concurrency

    Send the trace at the load you expect at peak and at a quiet hour. Record time to first token, inter-token latency, queue time and error rate from each engine’s metrics endpoint.

  4. 4. Test the constrained path

    Run your JSON schema or tool-call format through each structured-output backend and count invalid outputs.

  5. 5. Check operations

    Kill a node, upgrade a minor version and read the logs. The engine you can run at 03:00 matters more than a small difference in speed.

What should I check before I trust a vLLM or SGLang comparison chart?

  • The versions. Both projects release often, so a chart for last year’s release says little about today’s. Look for the engine versions and the date.
  • The model and precision. A result for one model at one quantisation does not carry to another. Mixture-of-experts and dense models behave differently.
  • The prompt mix. Prefix-caching benefits depend on how much of each prompt repeats. A chart built on unique prompts cannot show them, and one built on identical prompts overstates them.
  • The latency target. Throughput at unbounded latency is not the throughput you can sell. Ask for time to first token and inter-token latency at a stated concurrency.
  • The settings. Check whether each engine was tuned, and whether speculative decoding or structured output was on for one engine and off for the other.
  • Who ran it. A project, a vendor and an independent tester have different incentives. Prefer a method you can rerun.

Where Swfte fits

Swfte is not a serving engine, and you do not need Swfte to run vLLM or SGLang. Start with the engine’s own docs and a spare GPU.

Swfte Connect is the gateway in front of whichever engine you pick. It is one OpenAI-compatible API with routing, fallback chains, budgets and an audit event stream, and a model that is not in its catalogue is treated as bring-your-own-key with an optional base URL, so an OpenAI-compatible endpoint can sit behind it. Its free dashboard also works with vLLM or Ollama without routing prompts through the Swfte gateway.

The enterprise full-stack generator in Swfte’s self-deployment tooling bundles a vLLM container. It is available on request and needs a licence key. Swfte tooling does not claim SGLang support: if you choose SGLang, you operate it yourself and Connect can front it as any other OpenAI-compatible endpoint. See Connect self-deploy for what is generated, and deploy an open-source LLM for the wider path from model choice to monitoring.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

Is SGLang faster than vLLM?

Neither project’s documentation answers that for your workload, and this page publishes no benchmark. The SGLang paper reports gains over the systems of 2023, and the vLLM paper reports its own, but both engines have changed since. Replay your real prompts through both on your hardware, and compare time to first token, inter-token latency and queue time.

Do SGLang and vLLM both support the OpenAI API?

Yes, both document OpenAI-compatible servers. vLLM lists chat completions, completions, responses, embeddings and audio endpoints, and also an Anthropic Messages API. SGLang’s OpenAI-compatible section covers completions, vision and embeddings, and its structured-output page shows response_format with json_schema. Test the exact parameters your client sends, such as tool calls.

Which licence do SGLang and vLLM use?

Both repositories state Apache 2.0. That licence covers the engine code only. The model weights you serve carry their own licence, which can differ sharply, so read the model card separately. Apache 2.0 is a permissive licence, but have your own counsel read the text for your use.

Can I run SGLang or vLLM on AMD or non-NVIDIA hardware?

Both document support beyond NVIDIA. vLLM names AMD and Intel GPUs, several CPU architectures and plugins for TPU, Gaudi and Ascend. SGLang names AMD Instinct, Google TPU, Intel, Apple Silicon, Ascend and Moore Threads. Support varies by model and version, so check your accelerator on the project’s install page.

Does either engine need Swfte?

No. Both run on their own with an OpenAI-compatible endpoint. Swfte Connect is optional: it sits in front of the endpoint to add routing, fallbacks, budgets and audit events. Swfte’s generated enterprise bundle includes a vLLM container and is available on request, and SGLang is not part of that tooling.

Where do I find metrics for each engine?

vLLM lists a Prometheus-compatible /metrics endpoint, with metrics such as time to first token and KV-cache usage, and a /health endpoint. SGLang exposes /metrics when you start it with --enable-metrics and ships an example Grafana dashboard. Scrape both into the same dashboard so that a comparison uses one set of definitions.

Put one governed gateway in front of whichever engine you choose.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.