Engineering

Engineering: GPU infrastructure and LLM serving

Practical guides for people who run models on their own GPUs: what to optimise first, how to size memory, how to split a large model across devices, and what breaks in a cluster. Every figure links to a paper, a vendor page or an engine document.

Updated 6 October 2026

Start here

The four core guides

These four are written to be read together. The first covers kernels and tuning, the second the cluster, the third how to split a very large model, and the fourth the mistakes that waste the most GPU time.

Inference engines

Serving and tuning

Engines, schedulers and the evaluation work that comes before production.

Sizing

What a model needs to run

Memory and hardware requirements for specific open-weight models.

Infrastructure and control

Deploying on infrastructure you control

Procurement, isolation, regions and architecture for private and sovereign deployments.

Reference

Sizing formulas the guides share

The same four rules of thumb appear across the guides. Each links to the post that derives it and cites the source for its inputs.

QuantityFormulaNoteDerived in
Weightsparameters x bytes per parameterA 405B model is about 810 GB in BF16, 405 GB in FP8. Mixture-of-experts models need memory for total parameters, not active ones.Read the derivation
KV cache per token2 x layers x KV heads x head dimension x bytesLlama 3 70B in BF16: 80 layers, 8 KV heads of 128 dimensions, so 320 KiB per token.Read the derivation
Training state16 bytes per parameter for mixed-precision AdamFrom the ZeRO paper: 2 for weights, 2 for gradients, 12 for FP32 weights and optimiser states, before activations.Read the derivation
Decode ridge pointdense peak FLOP/s divided by memory bandwidthFor H100 SXM, about 989 TFLOPS dense BF16 over 3.35 TB/s is roughly 295 FLOPs per byte, so decode is memory-bound until the batch is in the hundreds.Read the derivation
Arithmetic from published model configurations and vendor datasheets. Dense tensor-core figures are half of the with-sparsity figures NVIDIA lists.

Common questions

What should I optimise first on a slow LLM endpoint?
Profile before changing anything. Check that the right attention backend is active, that the KV cache is sized for your real context lengths, and that batching and chunked prefill are tuned. Then look for gaps between kernels in a timeline, which point to launch overhead or host problems. Weight quantisation, speculative decoding and custom kernels come later.
How do I estimate how many GPUs a large model needs?
Start from memory: weights are parameters times bytes per parameter, then add KV cache for your context length and concurrency, plus engine overhead. A 405B dense model is about 405 GB in FP8, which fits in one 8-GPU node of 80 GB GPUs with roughly 235 GB left over. Throughput targets usually need more than the memory minimum.
When should I use tensor, pipeline or expert parallelism?
Use tensor parallelism inside one NVLink domain because it needs frequent all-reduces. Use pipeline parallelism to span nodes, since it sends smaller messages between stages. Use expert parallelism for large mixture-of-experts models with enough GPUs and an RDMA network, since it needs an all-to-all per MoE layer.
Do these guides contain Swfte benchmarks?
No. The guides quote figures only from cited papers, vendor pages and engine documentation, each with its hardware and date, and label our own arithmetic as arithmetic. They do not describe a specific Swfte cluster or publish Swfte performance numbers.
Where can I run open-weight models on infrastructure I control?
Swfte describes the options on its Deploy models, dedicated cloud and GPU pages, and the Connect model gateway can be self-deployed. The engineering guides here are general and apply whichever hosting route you choose.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.