Engineering: GPU infrastructure and LLM serving
Practical guides for people who run models on their own GPUs: what to optimise first, how to size memory, how to split a large model across devices, and what breaks in a cluster. Every figure links to a paper, a vendor page or an engine document.
Updated 6 October 2026
The four core guides
These four are written to be read together. The first covers kernels and tuning, the second the cluster, the third how to split a very large model, and the fourth the mistakes that waste the most GPU time.
- Kernel-Level Optimisations for LLM Inference and Training
FlashAttention, paged KV cache, continuous batching, speculative decoding, quantisation kernels, CUDA graphs, NCCL and Linux host tuning, with what to try first and how to measure it.
- Building a GPU Cluster for Large Models, and What Breaks
Memory maths for 70B, 400B+ and MoE models, interconnect and topology, storage, failures and stragglers, schedulers, power, EU constraints, a pre-flight checklist and a failure-mode table.
- Serving 400B+ Models: Tensor, Pipeline, Expert Parallelism
What each parallelism strategy sends over the wire, where it belongs, and how DeepSeek-V3 and SGLang deployments are laid out.
- Why the Cluster Is Rarely the Bottleneck: 10 Mistakes
Ten inference mistakes that waste GPU money, each with its symptom, fix and the metric that shows it.
Serving and tuning
Engines, schedulers and the evaluation work that comes before production.
- Continuous Batching for LLM Inference
How iteration-level scheduling works and where it falls short.
- vLLM Continuous Batching Deep Dive
PagedAttention, the scheduler and the tuning flags that matter.
- LLM Serving Frameworks 2026 Compared
vLLM, TGI, SGLang and TensorRT-LLM side by side.
- Batch Inference for LLMs
Patterns for moving non-interactive work off the latency path.
- Self-Hosted LLM Stack: vLLM, Ollama, LiteLLM and More
The layers of a self-hosted stack and the open-source tools for each.
- Evaluate an Open-Source LLM Before Production
A checklist covering licence, supply chain and task evaluations.
What a model needs to run
Memory and hardware requirements for specific open-weight models.
- DeepSeek-V4.1-Flash: the VRAM maths
Architecture and serving arithmetic for a very large open-weight model.
- Kimi K3 and Inkling: what it takes to run them
Power, GPU and memory reality for very large open-weight models.
Deploying on infrastructure you control
Procurement, isolation, regions and architecture for private and sovereign deployments.
- The GPU Problem: Securing AI Hardware
Vendor relationships and procurement pitfalls.
- Deploy an Open-Source LLM in the EU, Step by Step
Pinning revisions, verifying weights and serving with vLLM on EU infrastructure.
- Sovereign Inference in the EU: What In-Region Must Mean
Defining the boundary of in-region inference and a ten-question checklist.
- Air-Gapped LLM Deployment: A Practical Checklist
Staging weights, crossing the gap with verification and removing hidden egress.
- Confidential Computing for LLMs
What GPU trusted execution environments protect and what they cost.
- The AI DMZ
Controlled ingress, auditable execution and zero-trust egress for model deployments.
- A Sovereign AI Reference Architecture
Control and data planes, trust boundaries and the life of an agent action.
- Building a Sovereign AI Stack, Layer by Layer
The decisions to make at each of six layers.
Sizing formulas the guides share
The same four rules of thumb appear across the guides. Each links to the post that derives it and cites the source for its inputs.
| Quantity | Formula | Note | Derived in |
|---|---|---|---|
| Weights | parameters x bytes per parameter | A 405B model is about 810 GB in BF16, 405 GB in FP8. Mixture-of-experts models need memory for total parameters, not active ones. | Read the derivation |
| KV cache per token | 2 x layers x KV heads x head dimension x bytes | Llama 3 70B in BF16: 80 layers, 8 KV heads of 128 dimensions, so 320 KiB per token. | Read the derivation |
| Training state | 16 bytes per parameter for mixed-precision Adam | From the ZeRO paper: 2 for weights, 2 for gradients, 12 for FP32 weights and optimiser states, before activations. | Read the derivation |
| Decode ridge point | dense peak FLOP/s divided by memory bandwidth | For H100 SXM, about 989 TFLOPS dense BF16 over 3.35 TB/s is roughly 295 FLOPs per byte, so decode is memory-bound until the batch is in the hundreds. | Read the derivation |
Run it on your own infrastructure
The guides above are general and describe no specific Swfte cluster. If you want open-weight models on infrastructure you control, these pages describe the options.
Swfte pages
- Deploy modelsRun open-weight models on infrastructure you control
- Dedicated cloudDedicated, region-pinned deployments
- GPUInference and training capacity
- ConnectThe model gateway
- Self-deploy ConnectRun the gateway in your own environment
- Sovereign infrastructureLayer one of the platform
- Best sovereign cloud providersHosting options compared
Common questions
- What should I optimise first on a slow LLM endpoint?
- Profile before changing anything. Check that the right attention backend is active, that the KV cache is sized for your real context lengths, and that batching and chunked prefill are tuned. Then look for gaps between kernels in a timeline, which point to launch overhead or host problems. Weight quantisation, speculative decoding and custom kernels come later.
- How do I estimate how many GPUs a large model needs?
- Start from memory: weights are parameters times bytes per parameter, then add KV cache for your context length and concurrency, plus engine overhead. A 405B dense model is about 405 GB in FP8, which fits in one 8-GPU node of 80 GB GPUs with roughly 235 GB left over. Throughput targets usually need more than the memory minimum.
- When should I use tensor, pipeline or expert parallelism?
- Use tensor parallelism inside one NVLink domain because it needs frequent all-reduces. Use pipeline parallelism to span nodes, since it sends smaller messages between stages. Use expert parallelism for large mixture-of-experts models with enough GPUs and an RDMA network, since it needs an all-to-all per MoE layer.
- Do these guides contain Swfte benchmarks?
- No. The guides quote figures only from cited papers, vendor pages and engine documentation, each with its hardware and date, and label our own arithmetic as arithmetic. They do not describe a specific Swfte cluster or publish Swfte performance numbers.
- Where can I run open-weight models on infrastructure I control?
- Swfte describes the options on its Deploy models, dedicated cloud and GPU pages, and the Connect model gateway can be self-deployed. The engineering guides here are general and apply whichever hosting route you choose.