Kernel-Level Optimisations for LLM Inference and Training
What to try first for faster LLM inference and training, what each technique costs and how to measure it.
Most "my model is slow" tickets are not a kernel problem. They are a batching problem, a memory problem or a launch-overhead problem that shows up as slow kernels. Still, once the obvious things are fixed, what is left is kernel-level: how attention reads the KV cache, whether weights are streamed at 16 bits or 4, whether the CPU can keep the GPU fed, and how collectives overlap with compute.
This guide walks that list in the order we would try it: what each item does, what it costs, how to measure it and where it bites. Every number comes from a cited paper, vendor page or engine doc, with its hardware and date. Our own arithmetic is labelled as arithmetic, and a claim with no link is a judgement.
Companion reads: building a GPU cluster for large models, serving 400B+ models with tensor, pipeline and expert parallelism and the ten inference mistakes that cost the most.
Start with the bound: memory or compute
Every kernel is limited by one of two things: how fast it can move bytes, or how fast it can do arithmetic. The roofline model (Williams, Waterman and Patterson; Berkeley technical report 2008, CACM April 2009) ties these together with one number, arithmetic intensity: FLOPs performed per byte moved. Below the machine's ridge point you are memory-bound, above it compute-bound.
For an H100 SXM, NVIDIA's product page lists 3.35 TB/s of memory bandwidth and tensor-core figures quoted "with sparsity": 1,979 TFLOPS for FP16/BF16 and 3,958 for FP8. Dense figures are half of those, so roughly 989 TFLOPS BF16 and 1,979 TFLOPS FP8. Dividing (our arithmetic, not a vendor claim):
ridge point, BF16 = 989e12 FLOP/s / 3.35e12 B/s ~ 295 FLOP/byte
ridge point, FP8 = 1979e12 / 3.35e12 ~ 590 FLOP/byte
Now look at one decode step. Each BF16 weight is two bytes and takes part in two FLOPs per token in the batch. Intensity is therefore about the batch size B in FLOP/byte. With FP8 weights it is about 2B FLOP/byte against a ridge of 590, which again means B around 300. Ignoring attention and activations, a dense model on an H100 is memory-bound in decode until you batch roughly 300 sequences. Prefill, which processes thousands of tokens at once, is compute-bound.
That one fact explains most of the list below.
| Phase | Usually bound by | What helps | What does not |
|---|---|---|---|
| Decode, small batch | Weight and KV-cache reads (HBM bandwidth) | Lower-precision weights, bigger batches, speculative decoding, GQA/MLA, FP8 KV cache | A faster tensor core |
| Decode, large batch | KV-cache reads, then compute | Paged KV, FP8 KV, better attention kernels | More weight compression alone |
| Prefill | Tensor-core throughput, attention | FlashAttention-3 class kernels, FP8 matmuls, chunking | Weight-only INT4 (it dequantises into a compute-bound kernel) |
| Training step | Mix: matmul-bound inside layers, comms-bound between them | Fusion, overlap, FlashAttention, parallelism layout | Anything that ignores the network |
| Multi-GPU decode | Collective latency, not bandwidth | NVLink/NVSwitch, fewer and fatter collectives, CUDA graphs | Bigger links alone |
A worked floor, again as arithmetic: a 70B model in BF16 is 140 GB of weights. Across two H100s that is 6.7 TB/s of combined bandwidth, so one decode step cannot take less than about 21 ms at small batch, or about 48 tokens per second per stream. If you measure 15 tokens per second per stream and the arithmetic says 48, there is 3x to find. If you measure 45, stop tuning kernels.
Measure before you touch anything
Three tools, used in this order.
The PyTorch profiler. torch.profiler records CPU operators and CUDA kernels (ProfilerActivity.CPU, ProfilerActivity.CUDA), with schedule(wait, warmup, active) so you skip startup noise, record_shapes and profile_memory for shapes and allocations, and export_chrome_trace for a timeline you can open in Perfetto. Use it to answer one question: which operators dominate GPU time, and is the GPU ever idle?
Nsight Systems. Nsight Systems gives a system-wide timeline of CUDA API calls, kernels, NVTX ranges and NCCL collectives. nsys profile --trace=cuda,nvtx is the usual starting point. This is where you see gaps: the GPU idle while the CPU prepares the next launch, or a collective that starts late because one rank was slow.
Nsight Compute. Nsight Compute goes inside a single kernel. Its Speed of Light section reports achieved throughput of compute and memory units as a percentage of the theoretical maximum, and it can draw a roofline for the kernel. Use it only on the two or three kernels the earlier tools flagged, because it replays kernels and is slow.
A timeline with launch overhead looks like this:
GPU |kern| |kern| |kern| |kern| |kern| |kern|
CPU |launch..|launch..|launch..|launch..|launch..|
^ idle gaps: the GPU is waiting on the CPU
with a CUDA graph the same work becomes one launch:
GPU |kern|kern|kern|kern|kern|kern|
CPU |launch|
A protocol that keeps you honest: fix the workload (model, precision, your own prompt and output length mix, concurrency); record time-to-first-token, inter-token latency at p50 and p99, and throughput at a stated concurrency; warm up and repeat, because clocks, power limits and neighbours move numbers; change one thing per run and log the commit, driver, CUDA, engine version and flags; and check output quality for anything that changes numerics (quantisation, FP8 KV cache, different attention kernels, sampling with speculative decoding).
Attention: the FlashAttention family
Attention's naive implementation writes an N by N score matrix to HBM, reads it back for the softmax, writes it again, and reads it a third time for the multiplication with V. The cost is dominated by that traffic, not by the arithmetic.
FlashAttention (Dao, Fu, Ermon, Rudra and Ré, May 2022) is "IO-aware": it tiles the computation so blocks of Q, K and V live in on-chip SRAM and the full score matrix is never materialised in HBM. The result is exact attention, not an approximation, with memory linear in sequence length. The paper's own abstract-level results include a 3x speedup on GPT-2 at 1K tokens.
FlashAttention-2 (July 2023) cut non-matmul FLOPs and improved parallelism and work partitioning, reaching 50 to 73 percent of theoretical FLOPs/s on an A100 per the paper.
FlashAttention-3 (July 2024) targets Hopper: warp specialisation to overlap data movement and compute using TMA and tensor cores, interleaved matmul and softmax, and FP8 with block quantisation. The abstract reports 1.5 to 2.0 times the speed of FlashAttention-2 in FP16, reaching 740 TFLOPs/s on an H100 (75 percent utilisation), and about 1.2 PFLOPs/s in FP8.
Two practical points sit behind those numbers.
First, decode is a different shape. During decode the query length is 1, so parallelising over query blocks leaves most of the GPU idle at small batch and long context. Flash-Decoding (Stanford CRFM, October 2023) adds a split along the key/value length, runs the pieces in parallel and combines them with a log-sum-exp reduction. Modern engines ship decode-specific attention kernels for this reason. If your workload is long-context at low concurrency, check that your engine picks one.
Second, you rarely call these yourself. vLLM, SGLang, TensorRT-LLM and PyTorch's scaled_dot_product_attention select a backend. Your job is to confirm the right backend is active for your GPU generation (Hopper-specific kernels do nothing on Ampere) and that nothing forces a fallback, such as an unsupported head size, an unusual mask or a custom positional scheme.
Attention variants that shrink the KV cache. Grouped-query attention (Ainslie et al., 2023) uses fewer key/value heads than query heads; Llama 3 uses 8 at every size from 8B to 405B (paper). DeepSeek-V3 caches a 512-dimensional compressed latent plus a 64-dimensional rotary key per layer (report). These are model choices you cannot change later, but they should influence which model you deploy.
The KV cache: paging, prefix reuse and FP8
KV-cache memory per token is 2 (K and V) times layers times KV heads times head dimension times bytes per element. Our arithmetic from published configurations:
| Model | Layers | KV heads x head dim | KV cache per token (BF16) | At 128K tokens |
|---|---|---|---|---|
| Llama 3 70B | 80 | 8 x 128 | 320 KiB | about 40 GiB per sequence |
| Llama 3 405B | 126 | 8 x 128 | 504 KiB | about 63 GiB per sequence |
| DeepSeek-V3 (MLA) | 61 | (512 + 64) latent | about 69 KiB | about 8.6 GiB per sequence |
Layer and head counts are from the Llama 3 and DeepSeek-V3 papers; head dimension 128 follows from model dimension divided by 64 or 128 heads. At long context the cache, not the weights, decides how many sequences fit.
PagedAttention. Before PagedAttention (Kwon et al., SOSP 2023), serving systems reserved contiguous KV memory per request for the maximum length, so fragmentation and over-reservation capped batch size. PagedAttention splits the cache into fixed-size blocks, keeps a block table per sequence like an OS page table, and allocates on demand. The paper reports 2 to 4 times the throughput of FasterTransformer and Orca at the same latency, with bigger gains for longer sequences and larger models. Note the baselines are from 2023. Against a modern baseline the headline gain is smaller, because everyone now pages. vLLM's design page says plainly that its kernel walkthrough is historical and no longer describes the current code, so read it for the idea, not the implementation.
Prefix caching. If many requests share a prefix (a system prompt, a document, a chat history), the KV for that prefix can be reused. vLLM's automatic prefix caching reuses blocks when a new request shares a prefix with an earlier one; its doc is explicit that this only shortens prefill, not generation, and does nothing if prefixes do not repeat. SGLang's RadixAttention (paper, December 2023) organises the cache as a radix tree for the same purpose and reports up to 6.4 times higher throughput than the systems it compared against on multi-call programs and structured output. Treat that as an upper bound for prefix-heavy agent and RAG traffic.
FP8 KV cache. vLLM supports --kv-cache-dtype fp8, which roughly halves cache size against BF16 and lets more tokens fit. Its docs recommend calibrating scales with llm-compressor, note that per-attention-head scales need the Flash Attention backend, and say sliding-window layers are more sensitive. Run your own long-context evals before shipping it.
What to watch. Preemption: when the cache fills, the scheduler evicts and later recomputes sequences, and tail latency spikes. vLLM's optimisation guide lists the fixes (raise gpu_memory_utilization, lower max_num_seqs or max_num_batched_tokens, add parallelism); alert on the preemption counter.
Scheduling decides the shapes your kernels see
Kernels are only as good as the batches they receive.
Continuous batching. Orca (OSDI 2022) introduced iteration-level scheduling: the scheduler runs one forward pass at a time and can add or remove requests between passes, instead of holding a batch until its slowest member finishes. The paper reports 36.9 times the throughput of NVIDIA FasterTransformer at the same latency on a GPT-3 175B model. Today's engines all do this. Our continuous batching explainer and vLLM deep dive cover the mechanics.
Chunked prefill. A long prefill in the same batch as decodes stalls every decode stream for its duration. Sarathi-Serve (March 2024) splits prefills into near-equal chunks and schedules them alongside decodes without stalling them. It reports 2.6 times the serving capacity of vLLM on Mistral-7B on one A100, 3.7 times on Yi-34B on two A100s, and 5.6 times on Falcon-180B with pipeline parallelism, under latency targets, against the vLLM of that time. vLLM's V1 engine enables chunked prefill by default where it can, and the doc describes the dial: smaller max_num_batched_tokens (for example 2048) favours inter-token latency, larger values favour time-to-first-token, and above 8192 is recommended for throughput on smaller models.
Disaggregation. Prefill is compute-bound and decode is memory-bound, so they can run on separate GPUs. DistServe (OSDI 2024) reports serving 7.4 times more requests or meeting a 12.6 times tighter latency target than the systems it compared against. It adds a KV transfer between pools, so it needs a fast network and pays off at scale. Do chunked prefill first.
Fusion, CUDA graphs and torch.compile
Fusion. A transformer layer is a few big matmuls separated by many small element-wise operations: residual add, RMSNorm, rotary embedding, activation, gating. Run separately, each reads its input from HBM and writes its output back. Fusing them into one kernel keeps intermediates in registers. FlashAttention is the famous case; the same idea applies to fused RMSNorm plus residual, fused SwiGLU, and fusing QKV projections into one matmul. torch.compile does a lot of this automatically by tracing your model and generating fused kernels. PyTorch's tutorial attributes its speedups to less Python overhead and fewer GPU reads and writes. Graph breaks, where tracing gives up, forfeit the optimisation silently, so count them.
CUDA graphs. A decode step launches hundreds of small kernels. Each launch costs CPU time, and when kernels run for microseconds the GPU waits on the CPU. A CUDA graph records a sequence of operations once with stream capture (cudaStreamBeginCapture, cudaStreamEndCapture), instantiates it with cudaGraphInstantiate, and replays it with a single cudaGraphLaunch. NVIDIA's introduction explains the mechanism and shows the launch overhead disappearing for short kernels. This matters most for small-batch decode, which is exactly the latency-sensitive case.
The costs are real:
- Graphs are shape-specialised. vLLM's CUDA graph design doc says it caches graphs per batch descriptor and, when no match exists, runs in eager mode. Unusual batch shapes silently lose the benefit.
- Some operations cannot be captured. The same doc notes cascade attention is not CUDA-graph compatible, so it falls back to piecewise mode, where attention runs eagerly and the rest is captured.
- Graphs cost memory and startup time. The
FULL_AND_PIECEWISEmode "requires the most memory and takes the longest to capture," and--enforce-eagerskips capture for the fastest start at the cost of steady-state decode speed.
How to measure: compare inter-token latency at batch 1 to 8 with and without graphs, and look at the Nsight Systems timeline for CPU-side gaps between kernels.
Quantisation kernels
Quantisation is two separate things that people conflate: storing weights in fewer bits, and computing in fewer bits. The first cuts memory traffic and helps memory-bound decode. The second raises arithmetic throughput and helps compute-bound prefill and training.
| Scheme | Weights | Activations | Helps most | Main risk |
|---|---|---|---|---|
| BF16 / FP16 | 16 bit | 16 bit | Baseline | Memory |
| FP8 (W8A8) | 8 bit | 8 bit | Prefill and decode on Hopper, Ada and newer | Scale handling, outlier layers |
| INT8 W8A8 | 8 bit | 8 bit | Wide hardware support | Activation outliers need care |
| W4A16 (AWQ, GPTQ) | 4 bit | 16 bit | Small-batch decode | Little help for prefill, accuracy on hard tasks |
| FP8 KV cache | n/a | n/a | Long context, high concurrency | Quality on long-range tasks |
FP8. The FP8 formats paper (Micikevicius et al., September 2022) defines E4M3 (4-bit exponent, 3-bit mantissa) and E5M2 (5-bit exponent, 2-bit mantissa), and shows FP8 matching 16-bit quality across several architectures, including language models up to 175B parameters. In practice E4M3 is used for weights and activations in the forward pass. vLLM's quantisation page says FP8 W8A8 is supported on Ada, Hopper and AMD GPUs, so it is not available on A100.
AWQ. AWQ (Lin et al., MLSys 2024 best paper) observes that protecting about 1 percent of salient weight channels, identified from activation statistics, greatly reduces 4-bit quantisation error, and does so with a per-channel scaling rather than mixed precision, which keeps the kernel simple. The paper reports more than 3 times speedup over the Hugging Face FP16 implementation on desktop and mobile GPUs.
GPTQ. GPTQ (Frantar, Ashkboos, Hoefler and Alistarh, October 2022) quantises one-shot using approximate second-order information, handles 175B-parameter models in about four GPU hours at 3 to 4 bits, and reports end-to-end speedups of about 3.25 times on an A100 and 4.5 times on an A6000 against FP16.
Read those speedups as "weight-only quantisation against an unoptimised FP16 baseline on memory-bound decode". They are real and also not what you will see at batch 256 on a modern engine, where you are closer to compute-bound and dequantisation overhead eats into the gain. vLLM's compatibility table lists its Marlin kernels (GPTQ, AWQ, FP8 and FP4 weights) as requiring Turing or newer.
Before shipping a quantised model: calibrate on your domain, evaluate on your tasks (long outputs and tool use for agents; perplexity is a weak proxy), compare latency and throughput at your target concurrency rather than batch 1, and pin the checkpoint, library version and kernel backend, since kernels differ in numerics.
Speculative decoding
Decode is memory-bound, so verifying several tokens in one pass costs about the same as generating one. Speculative decoding exploits that: a cheap draft proposes k tokens, the big model scores all of them in one forward pass, and a rejection-sampling rule accepts a prefix. Leviathan, Kalman and Matias (ICML 2023) show this preserves the target model's output distribution exactly, with 2 to 3 times acceleration on T5-XXL. Chen et al. (DeepMind, February 2023) report 2 to 2.5 times on Chinchilla 70B in a distributed setup. EAGLE drafts at the feature level and reports 2.7 to 3.5 times latency speedup on LLaMA2-Chat 70B with the output distribution maintained.
Those are best cases measured at low batch sizes. As batch size rises, the verifier is no longer memory-bound and drafting becomes pure overhead. vLLM's speculative decoding page puts it in the same terms: model-based methods (EAGLE, MTP, draft models) give the best latency reduction, while n-gram and suffix methods give modest speedups without adding load at peak traffic, and gains differ between low and high request rates. It also warns that batch size changes can alter log-probabilities because of non-determinism in batched operations.
Try it for interactive, low-concurrency traffic with a good draft (an EAGLE or MTP head trained for your target). Skip it for saturated batch jobs. Measure acceptance rate; if it is low, the draft is not earning its keep.
Writing your own kernel
Write a custom kernel only after the profiler shows one operator dominating and no engine backend covers it. Then pick the lightest tool.
- Triton. Triton is a Python-based language and compiler for GPU kernels. It handles tiling and scheduling for you, and is the usual route for fused element-wise kernels and custom attention variants.
- CUTLASS and CuTe. CUTLASS is NVIDIA's collection of abstractions for high-performance GEMM, with CuTe for thread and data layouts. It covers FP8 (e4m3, e5m2), block-scaled FP4 and MXFP formats, Volta through Blackwell. Pick it for the last stretch of matmul performance or new low-precision formats.
- Raw CUDA. Only for what the other two cannot express.
Write a correctness test against a reference first, benchmark across the shapes you serve rather than one large square matrix, and budget for re-tuning on each GPU generation. Maintenance is the real price of a custom kernel.
Communication kernels: NCCL, overlap and expert all-to-all
Multi-GPU work adds a third thing to be bound by: the network.
Tensor parallelism shards each layer's matrices across GPUs and needs an all-reduce per layer (two in the forward pass of a Megatron-style transformer block, per the Megatron-LM paper). With 126 layers, as in Llama 3 405B, that is 252 collectives per token step (our arithmetic). The messages are small at decode, so latency, not bandwidth, dominates. NVIDIA's NVLink and NVSwitch inference post (August 2024) reports up to 1.5 times higher throughput for NVSwitch than point-to-point links on Llama 3.1 70B with H200 GPUs. Keep TP inside one NVLink domain. We expand on this in the 400B+ serving guide.
Pipeline parallelism sends activations between stages with point-to-point transfers and pays a bubble. Llama 3 uses an interleaved schedule where the bubble ratio is (PP-1)/(V x M), with V virtual stages per rank and M micro-batches (paper).
Expert parallelism shards MoE experts across GPUs and needs an all-to-all dispatch and combine per MoE layer. DeepEP is DeepSeek's open communication library with high-throughput and low-latency all-to-all kernels, FP8 dispatch and NVLink plus RDMA paths. The DeepSeek-V3 report describes custom cross-node all-to-all kernels that limit each token to at most 4 nodes and use few SMs for communication so compute can overlap. vLLM exposes backends through --all2all-backend (including deepep_high_throughput for prefill and deepep_low_latency for decode), and its expert-parallel doc warns that those two perform poorly for mixed workloads.
What to do about NCCL in practice:
- Run
nccl-tests(all_reduce_perf -b 8 -e 128M -f 2 -g 8) at the message sizes your parallelism produces. If bus bandwidth is far below what the hardware should give, fix topology and fabric before touching model code. - Set
NCCL_DEBUG=INFOfor one run and read the topology, transport and algorithm NCCL chose. The environment variable reference coversNCCL_SOCKET_IFNAME,NCCL_IB_HCA,NCCL_NET_GDR_LEVEL,NCCL_P2P_LEVEL,NCCL_IB_TIMEOUT,NCCL_ALGOandNCCL_PROTO. Treat overrides as a last resort: forcing an algorithm NCCL did not pick usually hides a topology problem. - In containers, GPUDirect RDMA must be enabled explicitly; vLLM's parallelism docs call this out.
- Overlap communication with compute where the framework allows it. Both DeepSeek-V3 and Llama 3 describe doing this on purpose.
Training-side kernels and utilisation
The training scoreboard is model FLOPs utilisation (MFU): achieved model FLOPs divided by peak. Reference points, each with its hardware:
- Llama 3 405B pre-training on up to 16K H100 GPUs: 38 to 43 percent BF16 MFU, per Table 4 of the paper (July 2024).
- MegaScale (ByteDance, February 2024): 55.2 percent MFU training a 175B model on 12,288 GPUs, a 1.34 times improvement over its Megatron-LM baseline.
- FlashAttention-2: 72 percent MFU for a single A100 end-to-end GPT-style training run, without the communication costs of scale.
The gap between the single-GPU figure and the big-cluster figures is communication, load imbalance and failures, not slow matmuls. What moves training MFU, in rough order of payoff:
- FlashAttention-class attention and fused kernels for the element-wise bulk.
- A parallelism layout that matches the network: Llama 3 orders dimensions as tensor, context, pipeline, data, with the chattiest innermost, "usually constrained to within the same server."
- Overlap of FSDP or ZeRO gathers and reduce-scatters with backward compute. ZeRO (Rajbhandari et al.) partitions optimiser state, gradients and parameters; the paper counts mixed-precision Adam at 16 bytes per parameter.
- Activation checkpointing only where memory forces it, because it adds recompute.
- FP8 matmuls where your framework supports them and your loss curves agree. The DeepSeek-V3 report describes fine-grained FP8 mixed-precision training at that scale; read its recipe before copying the idea.
- Fixing the straggler. One slow GPU slows every GPU: Llama 3's authors write that even a single straggler "can slow down thousands of other GPUs." ByteDance's straggler study (May 2025) finds stragglers can come from several causes, not only failed hardware.
Linux and host tuning
Host settings rarely change a kernel's speed. They fix the jitter around it: a GPU that waits on a CPU thread scheduled onto the wrong socket, a NIC interrupt landing on a busy core, a CPU that clocks down between launches. All of these appear as gaps in the Nsight Systems timeline.
| Setting | Why it matters | How to check | Risk and notes |
|---|---|---|---|
| NUMA binding | A process on socket 0 driving a GPU wired to socket 1 crosses the inter-socket link for every copy | nvidia-smi topo -m for GPU, NIC and CPU affinity; numastat counters numa_miss, other_node | Bind each rank to the CPUs and memory local to its GPU (numactl, or in Slurm --cpus-per-gpu and the Cores= field in gres.conf) |
| CPU governor | powersave or a slow-ramping governor lets cores clock down between kernel launches | Read scaling_governor per core | The performance governor requests the highest allowed frequency. With intel_pstate the kernel doc notes it "bypasses the scaling governor layer", so check your driver mode |
| IRQ affinity | NIC interrupts on cores that also run your launch threads add jitter | /proc/interrupts; set smp_affinity_list for NIC queues | Pin queues to cores local to the NIC. If irqbalance runs, confirm it is not rewriting your choices |
| Huge pages | Fewer TLB misses for large host-memory working sets (pinned buffers, offloaded KV, data loaders) | HugePages_Total, HugePages_Free in /proc/meminfo | Per the kernel doc, reserve with hugepagesz= and hugepages= at boot or vm.nr_hugepages at runtime. GPU kernels do not use host page tables, so expect effects only on CPU-side work. Measure; some systems prefer transparent huge pages off |
| GPUDirect RDMA | NIC reads and writes GPU memory directly, no bounce through host RAM | NCCL_DEBUG=INFO shows whether GDR is used; nvidia-peermem loaded | The GPUDirect RDMA docs: devices must share the same upstream PCIe root complex for good results, and it is incompatible with IOMMUs that translate addresses |
None of these is a first move. Do them after you have seen the gap in a timeline that they would close.
What to try first
In this order, because each step is cheaper and less risky than the next:
| Step | Change | Expected effect | Measure | Watch for |
|---|---|---|---|---|
| 1 | Use a current engine and the right attention backend for your GPU | Removes avoidable slow paths | Kernel names in the profiler | Silent fallbacks |
| 2 | Size the KV cache properly: memory utilisation, max sequences, FP8 KV if quality holds | More concurrent sequences, fewer preemptions | Preemption counter, p99 latency | Quality on long context |
| 3 | Tune batching: chunked prefill, max_num_batched_tokens | Better throughput/latency balance | TTFT, inter-token latency, throughput at fixed concurrency | Interactive p99 |
| 4 | Enable prefix caching if prefixes repeat | Cheaper prefill | Cache hit rate, TTFT | Zero gain without repetition |
| 5 | CUDA graphs and compile on for decode | Lower small-batch latency | Gaps in Nsight Systems | Eager fallbacks, memory |
| 6 | Weight quantisation (FP8 first on Hopper; W4A16 where memory forces it) | Fewer GPUs, faster decode | Quality evals plus latency at your concurrency | Accuracy regressions |
| 7 | Speculative decoding for low-concurrency latency | Lower latency per stream | Acceptance rate, latency at low and high load | Hurts saturated throughput |
| 8 | Parallelism layout and NCCL | Depends on topology | nccl-tests, collective time in Nsight | TP across slow links |
| 9 | Host tuning: NUMA, governor, IRQ | Less jitter | Timeline gaps, p99 | Hidden by earlier steps |
| 10 | Custom kernels | Only for a profiled hot operator | Nsight Compute roofline | Maintenance cost |
Pitfalls we would check first
- Benchmarking with one prompt length. Real traffic mixes lengths, and the mix changes every conclusion above.
- Reporting throughput without latency, or the reverse.
- Comparing against a two-year-old baseline. Paper speedups are against their era's baselines.
- Changing numerics and performance in the same run, so you cannot tell which broke quality.
- Reading a vendor figure quoted with sparsity as a dense figure.
- Tuning one GPU, then adding TP across a slower link and wondering why scaling is poor.
Where Swfte fits
This guide is general engineering advice and does not describe a specific Swfte cluster. If you want to run open-weight models on infrastructure you control, Deploy models and dedicated cloud describe the options, GPU covers compute, and Connect is the model gateway that you can self-deploy. The infrastructure layer page shows how these fit together, and best sovereign cloud providers compares hosting options. For the engine side, see the 2026 serving frameworks comparison and the self-hosted stack guide.
Frequently asked questions
Which optimisation should I try first for a slow LLM endpoint?
Check the basics (attention backend, KV cache size, batching and chunked prefill), then profile. The roofline arithmetic above gives a floor for decode speed; if you are near it, kernel tweaks will not help, and weight compression, batching or more GPUs will.
Does FlashAttention change model outputs?
FlashAttention computes exact attention, not an approximation, so results match the standard implementation up to floating-point differences from reordered arithmetic. Quantisation, FP8 KV cache and some speculative-decoding setups are different: they can change outputs and need evaluation.
Is INT4 always faster than FP8?
No. W4A16 cuts the bytes read per decode step, which helps at small batch. It still computes in 16-bit, so it does little for compute-bound prefill. On Hopper-class GPUs FP8 W8A8 is often the simpler first move. Measure at your own concurrency.
Do CUDA graphs help with long prompts?
Mostly not. They remove CPU launch overhead, which matters when kernels are tiny, as in small-batch decode. Prefill kernels are large and compute-bound, so launch overhead is a small share. Graphs also add memory cost and fall back to eager execution when the batch shape was not captured.
When is a custom Triton or CUTLASS kernel worth it?
When a profiler shows one operator dominating, no engine backend covers it, and you can write a reference test and benchmark across your real shapes. Otherwise the maintenance cost across GPU generations and engine versions outweighs the gain.
All links were read in October 2026. Engine flags and defaults change between releases, so check the version you run.