← The journal
Engineering

Serving 400B+ Models: Tensor, Pipeline, Expert Parallelism

How to split a 400B+ model across GPUs: where tensor, pipeline and expert parallelism fit, and what each costs.

Swfte Journal / Engineering

A 400B-plus model does not fit on one GPU, and often not on one node. Serving it means choosing how to cut it across devices, and every cut turns a memory problem into a communication problem. The choice between tensor, pipeline and expert parallelism is mostly a choice about which links your traffic crosses and how often.

This is the practical version: what each strategy sends over the wire, where it belongs, and what to measure. It builds on building a GPU cluster for large models and kernel-level optimisations, and pairs with the ten inference mistakes that cost the most.

First, does it fit?

Weights are parameters times bytes per parameter, and the arithmetic decides your minimum GPU count. With 8 H100 GPUs at 80 GB each (640 GB per the DGX H100 guide) or 8 H200 at 141 GB each (1,128 GB, our arithmetic from NVIDIA's page):

ModelWeights BF16Weights FP8One 8x H100 node (640 GB)One 8x H200 node (1,128 GB)
Llama 3.1 405B (dense)810 GB405 GBFP8 only, about 235 GB left for cacheBoth, with BF16 leaving about 318 GB
DeepSeek-V3 (671B total, 37B active)1.34 TB671 GBNoFP8 only, about 457 GB left

Cache is what is left, and it shrinks fast at long context: about 504 KiB per token for Llama 3 405B in BF16, so 63 GiB for one 128K-token sequence (derivation in the kernel guide). DeepSeek-V3's latent attention caches about 69 KiB per token. "Fits" has to include your real context lengths and concurrency, plus engine overhead and activation memory.

Tensor parallelism: split every layer

Tensor parallelism (TP) shards each layer's weight matrices across GPUs. Megatron-LM showed the pattern: each transformer block needs two all-reduces in the forward pass, one after attention and one after the MLP. For a 126-layer model like Llama 3 405B that is 252 all-reduces per token step (our arithmetic).

What this means in practice:

  • Latency-bound, not bandwidth-bound. At decode the messages are one activation vector per sequence, so each all-reduce is small and its latency dominates. Hundreds of them per step set a floor on token time.
  • It belongs inside the NVLink domain. NVIDIA's August 2024 analysis on H200 GPUs reports up to 1.5 times higher throughput on Llama 3.1 70B with NVSwitch than with point-to-point links, at the same GPU count. vLLM's docs say "efficient tensor parallelism requires fast internode communication, preferably through high-speed network adapters such as InfiniBand" if you must cross nodes, and recommend setting TP to the GPUs per node.
  • It cuts latency as well as memory. Each GPU reads only its shard of the weights, so aggregate memory bandwidth goes up with TP size. That is why TP is the default tool for latency, up to the point where all-reduce cost eats the gain.
  • Bigger TP is not free. Past one node, every layer pays a network hop twice. Llama 3's training setup puts tensor parallelism innermost and "usually constrained to within the same server."

Rule: TP as wide as one NVLink domain, no wider, unless you have measured otherwise.

Pipeline parallelism: split by layers

Pipeline parallelism (PP) assigns consecutive layers to different GPU groups and passes activations forward with point-to-point transfers. Traffic is small, one activation tensor per stage boundary, which is why PP tolerates slower links. vLLM's guidance is direct: if the model fits on one node, use TP; "if the model is too large for a single node, combine tensor parallelism with pipeline parallelism", with TP set to GPUs per node and PP to the number of nodes. It also helps with uneven GPU splits when there is no NVLink.

The costs:

  • Bubbles. A stage is idle until work reaches it. At training time the Llama 3 paper gives an interleaved-schedule bubble ratio of (PP-1)/(V x M), with V virtual stages per rank and M micro-batches. At inference the equivalent is keeping enough requests in flight to fill stages. Low concurrency makes PP wasteful.
  • Added latency. A token passes every stage in turn. PP increases capacity, not speed per token.
  • Imbalance. The first and last stages carry embeddings and output projection; Llama 3 removed a layer from each end stage to even things out.

Rule: use PP to get across nodes, after TP has filled each node. Keep it as shallow as possible.

Expert parallelism: split the experts

In a mixture-of-experts model, each token is routed to a few experts. Expert parallelism (EP) places different experts on different GPUs, so tokens must be sent to the GPUs holding their experts (dispatch) and the results sent back (combine). That is an all-to-all, twice per MoE layer.

Two things make EP different from TP:

  1. The batch per expert is small. Each expert only sees the tokens routed to it. DeepSeek's report says decode batches per expert are "usually within 256 tokens" and the bottleneck is memory access, not computation.
  2. Load is uneven. Popular experts become hotspots, and the slowest GPU gates the layer.

DeepSeek-V3's deployment is the best public reference, using different layouts per phase, on H800 GPUs with NVLink inside nodes and InfiniBand between them:

StageMinimum unitAttentionMoE
Prefill4 nodes, 32 GPUsTP4 with sequence parallelism, DP8EP32, with 32 redundant experts
Decode40 nodes, 320 GPUsTP4 with sequence parallelism, DP80EP320, one expert per GPU; 64 GPUs host redundant and shared experts

Notable details from that report: redundant copies of heavy experts are chosen from live load statistics and adjusted periodically (for example every 10 minutes); decode uses direct point-to-point IB transfers with IBGDA to minimise latency; two micro-batches overlap one's attention with the other's dispatch and combine; and the training-side all-to-all kernels cap each token to at most 4 nodes to limit cross-node IB traffic.

The open-source stack has caught up in pieces. DeepEP provides EP all-to-all kernels with FP8 dispatch and NVLink plus RDMA paths. vLLM's expert-parallel doc describes --enable-expert-parallel (EP size equals TP size times DP size), --enable-eplb for load balancing, and --all2all-backend options including deepep_high_throughput for prefill and deepep_low_latency for decode. It notes the high- and low-latency backends "show poor performance for mixed workloads" and that redundant experts cost memory (about 2.4 GB for one redundant expert per EP rank on DeepSeek-V3).

SGLang's team published a public reproduction on May 5, 2025: 12 nodes of 8 H100 GPUs (96 total), using prefill-decode disaggregation, DeepEP, EPLB and two-batch overlap, reporting 52.3k input and 22.3k output tokens per second per node for 2,000-token inputs (LMSYS blog). Those numbers are for that workload, that hardware and that date, and are not a forecast for yours.

Rule: EP pays off when the model is a large MoE, you have enough GPUs for each expert to see a decent batch, and you have an RDMA network and the engineering time to operate it. For a small MoE on one node, plain TP is usually simpler.

Choosing a layout

SituationStarting layoutWhy
Dense 400B, FP8, one 8-GPU node of 80 GBTP8Fits, all traffic on NVLink
Dense 400B, BF16, two H100 nodesTP8 x PP2TP stays on NVLink, only stage boundaries cross the network
Dense 400B on 8 x H200TP8Fits in BF16 or FP8 with more cache room
Large MoE (600B+), modest scaleTP8 x PP2 or TP across one node plus DPSimpler than EP; test it first
Large MoE, high throughput, many nodesDP attention plus EP MoE, EPLB, PD splitThe DeepSeek and SGLang pattern
Need more requests per second, not bigger modelAdd data-parallel replicasReplicas scale throughput without extra communication

Throughput comes from replicas, latency from TP. When one replica is already fast enough, adding replicas beats stretching a layout wider.

Operating a model that spans many GPUs

  • Blast radius. One failed GPU takes down the whole replica. A 320-GPU decode unit is a large blast radius, so plan health checks, fast restart and spares. Our cluster guide lists the Xid and ECC signals to watch.
  • Startup time. Loading 400 GB or more and capturing CUDA graphs takes minutes. vLLM notes that --enforce-eager starts fastest but loses steady-state decode speed, so you trade cold-start time against speed.
  • Preemption. If the cache fills, vLLM's guide lists the fixes: raise memory utilisation, lower max_num_seqs or max_num_batched_tokens, or add parallelism so each GPU has more cache room, each with a latency cost.
  • Disaggregation. DistServe reports serving 7.4 times more requests or meeting a 12.6 times tighter latency target by putting prefill and decode on separate GPUs. It needs a fast KV transfer path, so it suits large deployments.
  • Observability. Track per-expert token counts, all-to-all time, preemptions, TTFT and inter-token latency separately.

What to test before you commit

  1. nccl-tests at your TP and EP group sizes, on the real fabric.
  2. Decode latency at batch 1, 8 and your target concurrency for two or three layouts, with your real prompt-length mix.
  3. KV cache headroom and preemption rate at peak load.
  4. Behaviour when one rank dies: how long to detect, how long to recover.
  5. Quality of any quantised variant on your tasks, because FP8 versus BF16 is a numerics change.

Where Swfte fits

This is general guidance, not a description of a Swfte cluster. Deploy models and dedicated cloud describe running open-weight models on infrastructure you control, GPU covers compute, and Connect is the gateway you can self-deploy. See also the infrastructure layer and best sovereign cloud providers. For serving engines, read the 2026 serving frameworks comparison.

Frequently asked questions

What is the difference between tensor, pipeline and expert parallelism?

Tensor parallelism splits each layer's matrices across GPUs and needs frequent all-reduces, so it belongs inside one NVLink domain. Pipeline parallelism assigns groups of layers to different GPUs and passes activations between them, so it tolerates slower links. Expert parallelism places MoE experts on different GPUs and needs an all-to-all per MoE layer.

How many GPUs do I need for Llama 3.1 405B?

By memory, FP8 weights are about 405 GB, so one 8x H100 node (640 GB) is the minimum with around 235 GB left for cache and overhead. BF16 needs about 810 GB, so two H100 nodes or one 8x H200 node. Throughput and context targets usually require more.

Is pipeline parallelism bad for latency?

It adds latency per token because each token passes through every stage, and it needs enough concurrent requests to keep stages busy. Use it to span nodes after TP has filled each node, not as a way to go faster.

When should I use expert parallelism?

For large MoE models at enough scale that each expert gets a useful batch, on a fabric with RDMA, and when you can operate load balancing for hot experts. For smaller MoE models or modest traffic, tensor parallelism is simpler.

Does prefill and decode disaggregation help?

It can, at scale. DistServe reports large gains in requests served or latency targets met, and DeepSeek's deployment uses different layouts for prefill and decode. It adds a KV transfer between pools and more moving parts, so try chunked prefill and replicas first.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.