Serving engines
SGLang vs vLLM: which inference engine to run, by workload
Compare SGLang and vLLM on what each project documents, and pick one by workload instead of by a benchmark chart.
Both are open-source, Apache 2.0, OpenAI-compatible inference servers for GPUs you control, and both document prefix caching, speculative decoding, structured output, quantisation and multi-GPU parallelism. The documented differences are narrow: SGLang is built around RadixAttention and agentic workloads, vLLM around PagedAttention and the widest hardware and plugin list. This page publishes no throughput numbers, because the only way to know which is quicker for you is to replay your own traffic through both.
Last verified 2026-10-07. Sources are listed at the end of the page.
SGLang vs vLLM: what each project documents
Every cell below is taken from the project’s own documentation or repository on 2026-10-07. A blank in the project’s docs is written as “not stated”, never filled from memory.
| Topic | vLLM | SGLang |
|---|---|---|
| Licence | Apache-2.0 (repository). | Apache 2.0 (repository). The docs say the project is hosted under the non-profit open-source organisation LMSYS. |
| Core techniques named by the project | PagedAttention, continuous batching with chunked prefill and prefix caching, optimised attention kernels (FlashAttention, FlashInfer and others). | RadixAttention, prefix caching and multi-GPU parallelism (documentation home page). The SGLang paper describes RadixAttention as KV cache reuse. |
| Speculative decoding | n-gram, suffix, EAGLE and DFlash variants listed. | EAGLE-2 and EAGLE-3, MTP, UNO, DFLASH, a standalone draft model and an n-gram variant listed. |
| Structured output | xgrammar or guidance as backends; JSON schema, regex, grammar and choice; on by default in the OpenAI-compatible server. | XGrammar (default), Outlines or Llguidance; JSON schema, regex or EBNF, one constraint per request; usable through response_format. |
| Parallelism | Tensor, pipeline, data, expert and context parallelism. Multi-node through Ray or multiprocessing. | Tensor, pipeline, attention-context and MoE data parallelism, plus a multi-node flag for tensor parallelism. Data parallelism is launched through a router. |
| Quantisation | FP8, MXFP8 and MXFP4, NVFP4, INT8, INT4, GPTQ and AWQ, GGUF, compressed-tensors and more. | The server-arguments page lists awq, fp8, gptq, marlin, bitsandbytes and gguf, among others. |
| OpenAI-compatible API | Chat completions, completions, responses, embeddings, transcriptions and translations. Anthropic Messages API and gRPC are also listed. | Completions, vision and embeddings pages in the OpenAI-compatible section; the docs home page says it is compatible with Hugging Face and OpenAI APIs. |
| Hardware named | NVIDIA, AMD and Intel GPUs, x86, ARM and PowerPC CPUs, plus plugins for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend, Apple Silicon and others. | NVIDIA, AMD Instinct, Google TPU, Intel GPUs and CPUs, Apple Silicon, Huawei Ascend and Moore Threads, with specific cards listed in the README. |
| Disaggregated prefill and decode | Documented and marked experimental and subject to change. The page says it does not improve throughput. | Documented with Mooncake and NIXL transfer engines and single-node and multi-node examples. The page lists unsupported features. |
| Metrics and health | /metrics in Prometheus format and a /health endpoint, both listed. | /metrics when started with --enable-metrics, with time to first token, queue and token metrics. A health endpoint was not found on the pages read for this guide. |
Sources are listed at the foot of the page. Both projects release often, so re-check the version you will install.
What actually differs between them?
The two engines overlap more than their marketing suggests. vLLM’s documentation names automatic prefix caching, and SGLang’s names RadixAttention. Both reuse KV cache across requests that share a prefix. vLLM’s design page says it hashes each KV-cache block by the tokens in the block and the tokens before it. SGLang’s paper describes RadixAttention as KV cache reuse. Neither page tells you which wins on your prompts.
Where the projects present themselves differently is emphasis. vLLM describes itself as a library for LLM inference and serving and leads with memory management and breadth: hardware plugins, many quantisation formats, and an API server that lists more than chat completions. SGLang describes itself as an inference framework optimised for agentic workloads, reinforcement-learning rollouts and large-scale serving,.
The two papers each report large speedups, but against the baselines of their day, and both engines have moved on. This page does not quote them as a comparison. If a vendor chart ranks one engine, ask for the model, the GPU, the prompt mix, the concurrency, the engine version and the date.
Which should I pick for my workload?
These are decision rules drawn from the documented features. They tell you which engine to test first. They do not predict a winner.
| Workload | Test first | Why, and what to check |
|---|---|---|
| Shared multi-user chat on a common model | vLLM, then SGLang | Both serve the OpenAI API with continuous batching. Check that the exact model and quantisation you want is listed as supported by the version you install, and measure time to first token at your real concurrency. |
| Agent loops with long shared prefixes | SGLang, then vLLM | SGLang’s own pages lead with RadixAttention and agentic workloads. vLLM also has automatic prefix caching. Replay a real multi-turn trace through both and compare cache hit behaviour in their metrics. |
| Structured JSON output at volume | Either | Both document xgrammar. Differences are in the secondary backends (Outlines and Llguidance for SGLang, guidance for vLLM) and in schema features. Test your own JSON schema, including its hardest case, on both. |
| Hardware you already own | The one whose install page lists it | Both lists are long and differ at the edges: vLLM names Gaudi and IBM Spyre plugins, SGLang names Moore Threads. Check your exact accelerator and driver on each project’s install page before anything else. |
| Very large mixture-of-experts models across nodes | Whichever documents your model | Both document expert and tensor parallelism and multi-node launch. Model-specific recipes decide this, so read the model card’s serving section first. |
| Separating prefill and decode | SGLang, with care | SGLang documents it with two transfer engines. vLLM marks its version experimental. Both pages list limits, and neither is a reason to adopt the idea before you have a latency target it would fix. |
How do I compare them on my own traffic?
One afternoon of measurement is worth more than any published chart. Keep the setup identical for both engines.
1. Freeze the model artefact
Use the same weights and quantisation for both. If one engine needs a different format, record that and treat the pair as two separate candidates.
2. Capture real prompts
Export a few hundred anonymised requests with their real system prompts, tool definitions and conversation history. Prefix reuse only shows up when the prompts are real.
3. Replay at your concurrency
Send the trace at the load you expect at peak and at a quiet hour. Record time to first token, inter-token latency, queue time and error rate from each engine’s metrics endpoint.
4. Test the constrained path
Run your JSON schema or tool-call format through each structured-output backend and count invalid outputs.
5. Check operations
Kill a node, upgrade a minor version and read the logs. The engine you can run at 03:00 matters more than a small difference in speed.
What should I check before I trust a vLLM or SGLang comparison chart?
- The versions. Both projects release often, so a chart for last year’s release says little about today’s. Look for the engine versions and the date.
- The model and precision. A result for one model at one quantisation does not carry to another. Mixture-of-experts and dense models behave differently.
- The prompt mix. Prefix-caching benefits depend on how much of each prompt repeats. A chart built on unique prompts cannot show them, and one built on identical prompts overstates them.
- The latency target. Throughput at unbounded latency is not the throughput you can sell. Ask for time to first token and inter-token latency at a stated concurrency.
- The settings. Check whether each engine was tuned, and whether speculative decoding or structured output was on for one engine and off for the other.
- Who ran it. A project, a vendor and an independent tester have different incentives. Prefer a method you can rerun.
Where Swfte fits
Swfte is not a serving engine, and you do not need Swfte to run vLLM or SGLang. Start with the engine’s own docs and a spare GPU.
Swfte Connect is the gateway in front of whichever engine you pick. It is one OpenAI-compatible API with routing, fallback chains, budgets and an audit event stream, and a model that is not in its catalogue is treated as bring-your-own-key with an optional base URL, so an OpenAI-compatible endpoint can sit behind it. Its free dashboard also works with vLLM or Ollama without routing prompts through the Swfte gateway.
The enterprise full-stack generator in Swfte’s self-deployment tooling bundles a vLLM container. It is available on request and needs a licence key. Swfte tooling does not claim SGLang support: if you choose SGLang, you operate it yourself and Connect can front it as any other OpenAI-compatible endpoint. See Connect self-deploy for what is generated, and deploy an open-source LLM for the wider path from model choice to monitoring.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- vLLM documentation home. Core techniques, quantisation formats, parallelism and hardware as listed.
- vLLM repository. Licence and README feature and hardware list.
- vLLM OpenAI-compatible server. Endpoint list, /health and /metrics.
- vLLM structured outputs. Backends and constraint types.
- vLLM automatic prefix caching design. How KV-cache blocks are hashed.
- vLLM disaggregated prefilling. Experimental status and the throughput statement.
- vLLM multi-node and parallelism guidance. Ray and multiprocessing runtimes, tensor and pipeline parallel advice.
- vLLM metrics design. Prometheus metric names.
- SGLang documentation home. Core techniques, hardware and the LMSYS hosting statement.
- SGLang repository. Licence, README description and hardware list.
- SGLang structured outputs. Grammar backends and constraint types.
- SGLang prefill-decode disaggregation. Transfer engines and listed limits.
- SGLang speculative decoding. Methods listed.
- SGLang production metrics. /metrics flag and metric categories.
- SGLang server arguments. Parallelism and quantisation flags.
- SGLang OpenAI-compatible APIs. Which OpenAI endpoints the docs cover.
- SGLang paper (arXiv 2312.07104). RadixAttention described as KV cache reuse.
- PagedAttention paper (arXiv 2309.06180). PagedAttention definition.
Frequently asked questions
Is SGLang faster than vLLM?
Neither project’s documentation answers that for your workload, and this page publishes no benchmark. The SGLang paper reports gains over the systems of 2023, and the vLLM paper reports its own, but both engines have changed since. Replay your real prompts through both on your hardware, and compare time to first token, inter-token latency and queue time.
Do SGLang and vLLM both support the OpenAI API?
Yes, both document OpenAI-compatible servers. vLLM lists chat completions, completions, responses, embeddings and audio endpoints, and also an Anthropic Messages API. SGLang’s OpenAI-compatible section covers completions, vision and embeddings, and its structured-output page shows response_format with json_schema. Test the exact parameters your client sends, such as tool calls.
Which licence do SGLang and vLLM use?
Both repositories state Apache 2.0. That licence covers the engine code only. The model weights you serve carry their own licence, which can differ sharply, so read the model card separately. Apache 2.0 is a permissive licence, but have your own counsel read the text for your use.
Can I run SGLang or vLLM on AMD or non-NVIDIA hardware?
Both document support beyond NVIDIA. vLLM names AMD and Intel GPUs, several CPU architectures and plugins for TPU, Gaudi and Ascend. SGLang names AMD Instinct, Google TPU, Intel, Apple Silicon, Ascend and Moore Threads. Support varies by model and version, so check your accelerator on the project’s install page.
Does either engine need Swfte?
No. Both run on their own with an OpenAI-compatible endpoint. Swfte Connect is optional: it sits in front of the endpoint to add routing, fallbacks, budgets and audit events. Swfte’s generated enterprise bundle includes a vLLM container and is available on request, and SGLang is not part of that tooling.
Where do I find metrics for each engine?
vLLM lists a Prometheus-compatible /metrics endpoint, with metrics such as time to first token and KV-cache usage, and a /health endpoint. SGLang exposes /metrics when you start it with --enable-metrics and ships an example Grafana dashboard. Scrape both into the same dashboard so that a comparison uses one set of definitions.