← The journal
Engineering

Kernel-Level Optimisations for LLM Inference and Training

What to try first for faster LLM inference and training, what each technique costs and how to measure it.

Swfte Journal / Engineering

Most "my model is slow" tickets are not a kernel problem. They are a batching problem, a memory problem or a launch-overhead problem that shows up as slow kernels. Still, once the obvious things are fixed, what is left is kernel-level: how attention reads the KV cache, whether weights are streamed at 16 bits or 4, whether the CPU can keep the GPU fed, and how collectives overlap with compute.

This guide walks that list in the order we would try it: what each item does, what it costs, how to measure it and where it bites. Every number comes from a cited paper, vendor page or engine doc, with its hardware and date. Our own arithmetic is labelled as arithmetic, and a claim with no link is a judgement.

Companion reads: building a GPU cluster for large models, serving 400B+ models with tensor, pipeline and expert parallelism and the ten inference mistakes that cost the most.

Start with the bound: memory or compute

Every kernel is limited by one of two things: how fast it can move bytes, or how fast it can do arithmetic. The roofline model (Williams, Waterman and Patterson; Berkeley technical report 2008, CACM April 2009) ties these together with one number, arithmetic intensity: FLOPs performed per byte moved. Below the machine's ridge point you are memory-bound, above it compute-bound.

For an H100 SXM, NVIDIA's product page lists 3.35 TB/s of memory bandwidth and tensor-core figures quoted "with sparsity": 1,979 TFLOPS for FP16/BF16 and 3,958 for FP8. Dense figures are half of those, so roughly 989 TFLOPS BF16 and 1,979 TFLOPS FP8. Dividing (our arithmetic, not a vendor claim):

ridge point, BF16  = 989e12 FLOP/s / 3.35e12 B/s  ~ 295 FLOP/byte
ridge point, FP8   = 1979e12 / 3.35e12            ~ 590 FLOP/byte

Now look at one decode step. Each BF16 weight is two bytes and takes part in two FLOPs per token in the batch. Intensity is therefore about the batch size B in FLOP/byte. With FP8 weights it is about 2B FLOP/byte against a ridge of 590, which again means B around 300. Ignoring attention and activations, a dense model on an H100 is memory-bound in decode until you batch roughly 300 sequences. Prefill, which processes thousands of tokens at once, is compute-bound.

That one fact explains most of the list below.

PhaseUsually bound byWhat helpsWhat does not
Decode, small batchWeight and KV-cache reads (HBM bandwidth)Lower-precision weights, bigger batches, speculative decoding, GQA/MLA, FP8 KV cacheA faster tensor core
Decode, large batchKV-cache reads, then computePaged KV, FP8 KV, better attention kernelsMore weight compression alone
PrefillTensor-core throughput, attentionFlashAttention-3 class kernels, FP8 matmuls, chunkingWeight-only INT4 (it dequantises into a compute-bound kernel)
Training stepMix: matmul-bound inside layers, comms-bound between themFusion, overlap, FlashAttention, parallelism layoutAnything that ignores the network
Multi-GPU decodeCollective latency, not bandwidthNVLink/NVSwitch, fewer and fatter collectives, CUDA graphsBigger links alone

A worked floor, again as arithmetic: a 70B model in BF16 is 140 GB of weights. Across two H100s that is 6.7 TB/s of combined bandwidth, so one decode step cannot take less than about 21 ms at small batch, or about 48 tokens per second per stream. If you measure 15 tokens per second per stream and the arithmetic says 48, there is 3x to find. If you measure 45, stop tuning kernels.

Measure before you touch anything

Three tools, used in this order.

The PyTorch profiler. torch.profiler records CPU operators and CUDA kernels (ProfilerActivity.CPU, ProfilerActivity.CUDA), with schedule(wait, warmup, active) so you skip startup noise, record_shapes and profile_memory for shapes and allocations, and export_chrome_trace for a timeline you can open in Perfetto. Use it to answer one question: which operators dominate GPU time, and is the GPU ever idle?

Nsight Systems. Nsight Systems gives a system-wide timeline of CUDA API calls, kernels, NVTX ranges and NCCL collectives. nsys profile --trace=cuda,nvtx is the usual starting point. This is where you see gaps: the GPU idle while the CPU prepares the next launch, or a collective that starts late because one rank was slow.

Nsight Compute. Nsight Compute goes inside a single kernel. Its Speed of Light section reports achieved throughput of compute and memory units as a percentage of the theoretical maximum, and it can draw a roofline for the kernel. Use it only on the two or three kernels the earlier tools flagged, because it replays kernels and is slow.

A timeline with launch overhead looks like this:

GPU  |kern|  |kern|  |kern|  |kern|  |kern|  |kern|
CPU  |launch..|launch..|launch..|launch..|launch..|
          ^ idle gaps: the GPU is waiting on the CPU

with a CUDA graph the same work becomes one launch:
GPU  |kern|kern|kern|kern|kern|kern|
CPU  |launch|

A protocol that keeps you honest: fix the workload (model, precision, your own prompt and output length mix, concurrency); record time-to-first-token, inter-token latency at p50 and p99, and throughput at a stated concurrency; warm up and repeat, because clocks, power limits and neighbours move numbers; change one thing per run and log the commit, driver, CUDA, engine version and flags; and check output quality for anything that changes numerics (quantisation, FP8 KV cache, different attention kernels, sampling with speculative decoding).

Attention: the FlashAttention family

Attention's naive implementation writes an N by N score matrix to HBM, reads it back for the softmax, writes it again, and reads it a third time for the multiplication with V. The cost is dominated by that traffic, not by the arithmetic.

FlashAttention (Dao, Fu, Ermon, Rudra and Ré, May 2022) is "IO-aware": it tiles the computation so blocks of Q, K and V live in on-chip SRAM and the full score matrix is never materialised in HBM. The result is exact attention, not an approximation, with memory linear in sequence length. The paper's own abstract-level results include a 3x speedup on GPT-2 at 1K tokens.

FlashAttention-2 (July 2023) cut non-matmul FLOPs and improved parallelism and work partitioning, reaching 50 to 73 percent of theoretical FLOPs/s on an A100 per the paper.

FlashAttention-3 (July 2024) targets Hopper: warp specialisation to overlap data movement and compute using TMA and tensor cores, interleaved matmul and softmax, and FP8 with block quantisation. The abstract reports 1.5 to 2.0 times the speed of FlashAttention-2 in FP16, reaching 740 TFLOPs/s on an H100 (75 percent utilisation), and about 1.2 PFLOPs/s in FP8.

Two practical points sit behind those numbers.

First, decode is a different shape. During decode the query length is 1, so parallelising over query blocks leaves most of the GPU idle at small batch and long context. Flash-Decoding (Stanford CRFM, October 2023) adds a split along the key/value length, runs the pieces in parallel and combines them with a log-sum-exp reduction. Modern engines ship decode-specific attention kernels for this reason. If your workload is long-context at low concurrency, check that your engine picks one.

Second, you rarely call these yourself. vLLM, SGLang, TensorRT-LLM and PyTorch's scaled_dot_product_attention select a backend. Your job is to confirm the right backend is active for your GPU generation (Hopper-specific kernels do nothing on Ampere) and that nothing forces a fallback, such as an unsupported head size, an unusual mask or a custom positional scheme.

Attention variants that shrink the KV cache. Grouped-query attention (Ainslie et al., 2023) uses fewer key/value heads than query heads; Llama 3 uses 8 at every size from 8B to 405B (paper). DeepSeek-V3 caches a 512-dimensional compressed latent plus a 64-dimensional rotary key per layer (report). These are model choices you cannot change later, but they should influence which model you deploy.

The KV cache: paging, prefix reuse and FP8

KV-cache memory per token is 2 (K and V) times layers times KV heads times head dimension times bytes per element. Our arithmetic from published configurations:

ModelLayersKV heads x head dimKV cache per token (BF16)At 128K tokens
Llama 3 70B808 x 128320 KiBabout 40 GiB per sequence
Llama 3 405B1268 x 128504 KiBabout 63 GiB per sequence
DeepSeek-V3 (MLA)61(512 + 64) latentabout 69 KiBabout 8.6 GiB per sequence

Layer and head counts are from the Llama 3 and DeepSeek-V3 papers; head dimension 128 follows from model dimension divided by 64 or 128 heads. At long context the cache, not the weights, decides how many sequences fit.

PagedAttention. Before PagedAttention (Kwon et al., SOSP 2023), serving systems reserved contiguous KV memory per request for the maximum length, so fragmentation and over-reservation capped batch size. PagedAttention splits the cache into fixed-size blocks, keeps a block table per sequence like an OS page table, and allocates on demand. The paper reports 2 to 4 times the throughput of FasterTransformer and Orca at the same latency, with bigger gains for longer sequences and larger models. Note the baselines are from 2023. Against a modern baseline the headline gain is smaller, because everyone now pages. vLLM's design page says plainly that its kernel walkthrough is historical and no longer describes the current code, so read it for the idea, not the implementation.

Prefix caching. If many requests share a prefix (a system prompt, a document, a chat history), the KV for that prefix can be reused. vLLM's automatic prefix caching reuses blocks when a new request shares a prefix with an earlier one; its doc is explicit that this only shortens prefill, not generation, and does nothing if prefixes do not repeat. SGLang's RadixAttention (paper, December 2023) organises the cache as a radix tree for the same purpose and reports up to 6.4 times higher throughput than the systems it compared against on multi-call programs and structured output. Treat that as an upper bound for prefix-heavy agent and RAG traffic.

FP8 KV cache. vLLM supports --kv-cache-dtype fp8, which roughly halves cache size against BF16 and lets more tokens fit. Its docs recommend calibrating scales with llm-compressor, note that per-attention-head scales need the Flash Attention backend, and say sliding-window layers are more sensitive. Run your own long-context evals before shipping it.

What to watch. Preemption: when the cache fills, the scheduler evicts and later recomputes sequences, and tail latency spikes. vLLM's optimisation guide lists the fixes (raise gpu_memory_utilization, lower max_num_seqs or max_num_batched_tokens, add parallelism); alert on the preemption counter.

Scheduling decides the shapes your kernels see

Kernels are only as good as the batches they receive.

Continuous batching. Orca (OSDI 2022) introduced iteration-level scheduling: the scheduler runs one forward pass at a time and can add or remove requests between passes, instead of holding a batch until its slowest member finishes. The paper reports 36.9 times the throughput of NVIDIA FasterTransformer at the same latency on a GPT-3 175B model. Today's engines all do this. Our continuous batching explainer and vLLM deep dive cover the mechanics.

Chunked prefill. A long prefill in the same batch as decodes stalls every decode stream for its duration. Sarathi-Serve (March 2024) splits prefills into near-equal chunks and schedules them alongside decodes without stalling them. It reports 2.6 times the serving capacity of vLLM on Mistral-7B on one A100, 3.7 times on Yi-34B on two A100s, and 5.6 times on Falcon-180B with pipeline parallelism, under latency targets, against the vLLM of that time. vLLM's V1 engine enables chunked prefill by default where it can, and the doc describes the dial: smaller max_num_batched_tokens (for example 2048) favours inter-token latency, larger values favour time-to-first-token, and above 8192 is recommended for throughput on smaller models.

Disaggregation. Prefill is compute-bound and decode is memory-bound, so they can run on separate GPUs. DistServe (OSDI 2024) reports serving 7.4 times more requests or meeting a 12.6 times tighter latency target than the systems it compared against. It adds a KV transfer between pools, so it needs a fast network and pays off at scale. Do chunked prefill first.

Fusion, CUDA graphs and torch.compile

Fusion. A transformer layer is a few big matmuls separated by many small element-wise operations: residual add, RMSNorm, rotary embedding, activation, gating. Run separately, each reads its input from HBM and writes its output back. Fusing them into one kernel keeps intermediates in registers. FlashAttention is the famous case; the same idea applies to fused RMSNorm plus residual, fused SwiGLU, and fusing QKV projections into one matmul. torch.compile does a lot of this automatically by tracing your model and generating fused kernels. PyTorch's tutorial attributes its speedups to less Python overhead and fewer GPU reads and writes. Graph breaks, where tracing gives up, forfeit the optimisation silently, so count them.

CUDA graphs. A decode step launches hundreds of small kernels. Each launch costs CPU time, and when kernels run for microseconds the GPU waits on the CPU. A CUDA graph records a sequence of operations once with stream capture (cudaStreamBeginCapture, cudaStreamEndCapture), instantiates it with cudaGraphInstantiate, and replays it with a single cudaGraphLaunch. NVIDIA's introduction explains the mechanism and shows the launch overhead disappearing for short kernels. This matters most for small-batch decode, which is exactly the latency-sensitive case.

The costs are real:

  • Graphs are shape-specialised. vLLM's CUDA graph design doc says it caches graphs per batch descriptor and, when no match exists, runs in eager mode. Unusual batch shapes silently lose the benefit.
  • Some operations cannot be captured. The same doc notes cascade attention is not CUDA-graph compatible, so it falls back to piecewise mode, where attention runs eagerly and the rest is captured.
  • Graphs cost memory and startup time. The FULL_AND_PIECEWISE mode "requires the most memory and takes the longest to capture," and --enforce-eager skips capture for the fastest start at the cost of steady-state decode speed.

How to measure: compare inter-token latency at batch 1 to 8 with and without graphs, and look at the Nsight Systems timeline for CPU-side gaps between kernels.

Quantisation kernels

Quantisation is two separate things that people conflate: storing weights in fewer bits, and computing in fewer bits. The first cuts memory traffic and helps memory-bound decode. The second raises arithmetic throughput and helps compute-bound prefill and training.

SchemeWeightsActivationsHelps mostMain risk
BF16 / FP1616 bit16 bitBaselineMemory
FP8 (W8A8)8 bit8 bitPrefill and decode on Hopper, Ada and newerScale handling, outlier layers
INT8 W8A88 bit8 bitWide hardware supportActivation outliers need care
W4A16 (AWQ, GPTQ)4 bit16 bitSmall-batch decodeLittle help for prefill, accuracy on hard tasks
FP8 KV cachen/an/aLong context, high concurrencyQuality on long-range tasks

FP8. The FP8 formats paper (Micikevicius et al., September 2022) defines E4M3 (4-bit exponent, 3-bit mantissa) and E5M2 (5-bit exponent, 2-bit mantissa), and shows FP8 matching 16-bit quality across several architectures, including language models up to 175B parameters. In practice E4M3 is used for weights and activations in the forward pass. vLLM's quantisation page says FP8 W8A8 is supported on Ada, Hopper and AMD GPUs, so it is not available on A100.

AWQ. AWQ (Lin et al., MLSys 2024 best paper) observes that protecting about 1 percent of salient weight channels, identified from activation statistics, greatly reduces 4-bit quantisation error, and does so with a per-channel scaling rather than mixed precision, which keeps the kernel simple. The paper reports more than 3 times speedup over the Hugging Face FP16 implementation on desktop and mobile GPUs.

GPTQ. GPTQ (Frantar, Ashkboos, Hoefler and Alistarh, October 2022) quantises one-shot using approximate second-order information, handles 175B-parameter models in about four GPU hours at 3 to 4 bits, and reports end-to-end speedups of about 3.25 times on an A100 and 4.5 times on an A6000 against FP16.

Read those speedups as "weight-only quantisation against an unoptimised FP16 baseline on memory-bound decode". They are real and also not what you will see at batch 256 on a modern engine, where you are closer to compute-bound and dequantisation overhead eats into the gain. vLLM's compatibility table lists its Marlin kernels (GPTQ, AWQ, FP8 and FP4 weights) as requiring Turing or newer.

Before shipping a quantised model: calibrate on your domain, evaluate on your tasks (long outputs and tool use for agents; perplexity is a weak proxy), compare latency and throughput at your target concurrency rather than batch 1, and pin the checkpoint, library version and kernel backend, since kernels differ in numerics.

Speculative decoding

Decode is memory-bound, so verifying several tokens in one pass costs about the same as generating one. Speculative decoding exploits that: a cheap draft proposes k tokens, the big model scores all of them in one forward pass, and a rejection-sampling rule accepts a prefix. Leviathan, Kalman and Matias (ICML 2023) show this preserves the target model's output distribution exactly, with 2 to 3 times acceleration on T5-XXL. Chen et al. (DeepMind, February 2023) report 2 to 2.5 times on Chinchilla 70B in a distributed setup. EAGLE drafts at the feature level and reports 2.7 to 3.5 times latency speedup on LLaMA2-Chat 70B with the output distribution maintained.

Those are best cases measured at low batch sizes. As batch size rises, the verifier is no longer memory-bound and drafting becomes pure overhead. vLLM's speculative decoding page puts it in the same terms: model-based methods (EAGLE, MTP, draft models) give the best latency reduction, while n-gram and suffix methods give modest speedups without adding load at peak traffic, and gains differ between low and high request rates. It also warns that batch size changes can alter log-probabilities because of non-determinism in batched operations.

Try it for interactive, low-concurrency traffic with a good draft (an EAGLE or MTP head trained for your target). Skip it for saturated batch jobs. Measure acceptance rate; if it is low, the draft is not earning its keep.

Writing your own kernel

Write a custom kernel only after the profiler shows one operator dominating and no engine backend covers it. Then pick the lightest tool.

  • Triton. Triton is a Python-based language and compiler for GPU kernels. It handles tiling and scheduling for you, and is the usual route for fused element-wise kernels and custom attention variants.
  • CUTLASS and CuTe. CUTLASS is NVIDIA's collection of abstractions for high-performance GEMM, with CuTe for thread and data layouts. It covers FP8 (e4m3, e5m2), block-scaled FP4 and MXFP formats, Volta through Blackwell. Pick it for the last stretch of matmul performance or new low-precision formats.
  • Raw CUDA. Only for what the other two cannot express.

Write a correctness test against a reference first, benchmark across the shapes you serve rather than one large square matrix, and budget for re-tuning on each GPU generation. Maintenance is the real price of a custom kernel.

Communication kernels: NCCL, overlap and expert all-to-all

Multi-GPU work adds a third thing to be bound by: the network.

Tensor parallelism shards each layer's matrices across GPUs and needs an all-reduce per layer (two in the forward pass of a Megatron-style transformer block, per the Megatron-LM paper). With 126 layers, as in Llama 3 405B, that is 252 collectives per token step (our arithmetic). The messages are small at decode, so latency, not bandwidth, dominates. NVIDIA's NVLink and NVSwitch inference post (August 2024) reports up to 1.5 times higher throughput for NVSwitch than point-to-point links on Llama 3.1 70B with H200 GPUs. Keep TP inside one NVLink domain. We expand on this in the 400B+ serving guide.

Pipeline parallelism sends activations between stages with point-to-point transfers and pays a bubble. Llama 3 uses an interleaved schedule where the bubble ratio is (PP-1)/(V x M), with V virtual stages per rank and M micro-batches (paper).

Expert parallelism shards MoE experts across GPUs and needs an all-to-all dispatch and combine per MoE layer. DeepEP is DeepSeek's open communication library with high-throughput and low-latency all-to-all kernels, FP8 dispatch and NVLink plus RDMA paths. The DeepSeek-V3 report describes custom cross-node all-to-all kernels that limit each token to at most 4 nodes and use few SMs for communication so compute can overlap. vLLM exposes backends through --all2all-backend (including deepep_high_throughput for prefill and deepep_low_latency for decode), and its expert-parallel doc warns that those two perform poorly for mixed workloads.

What to do about NCCL in practice:

  • Run nccl-tests (all_reduce_perf -b 8 -e 128M -f 2 -g 8) at the message sizes your parallelism produces. If bus bandwidth is far below what the hardware should give, fix topology and fabric before touching model code.
  • Set NCCL_DEBUG=INFO for one run and read the topology, transport and algorithm NCCL chose. The environment variable reference covers NCCL_SOCKET_IFNAME, NCCL_IB_HCA, NCCL_NET_GDR_LEVEL, NCCL_P2P_LEVEL, NCCL_IB_TIMEOUT, NCCL_ALGO and NCCL_PROTO. Treat overrides as a last resort: forcing an algorithm NCCL did not pick usually hides a topology problem.
  • In containers, GPUDirect RDMA must be enabled explicitly; vLLM's parallelism docs call this out.
  • Overlap communication with compute where the framework allows it. Both DeepSeek-V3 and Llama 3 describe doing this on purpose.

Training-side kernels and utilisation

The training scoreboard is model FLOPs utilisation (MFU): achieved model FLOPs divided by peak. Reference points, each with its hardware:

  • Llama 3 405B pre-training on up to 16K H100 GPUs: 38 to 43 percent BF16 MFU, per Table 4 of the paper (July 2024).
  • MegaScale (ByteDance, February 2024): 55.2 percent MFU training a 175B model on 12,288 GPUs, a 1.34 times improvement over its Megatron-LM baseline.
  • FlashAttention-2: 72 percent MFU for a single A100 end-to-end GPT-style training run, without the communication costs of scale.

The gap between the single-GPU figure and the big-cluster figures is communication, load imbalance and failures, not slow matmuls. What moves training MFU, in rough order of payoff:

  1. FlashAttention-class attention and fused kernels for the element-wise bulk.
  2. A parallelism layout that matches the network: Llama 3 orders dimensions as tensor, context, pipeline, data, with the chattiest innermost, "usually constrained to within the same server."
  3. Overlap of FSDP or ZeRO gathers and reduce-scatters with backward compute. ZeRO (Rajbhandari et al.) partitions optimiser state, gradients and parameters; the paper counts mixed-precision Adam at 16 bytes per parameter.
  4. Activation checkpointing only where memory forces it, because it adds recompute.
  5. FP8 matmuls where your framework supports them and your loss curves agree. The DeepSeek-V3 report describes fine-grained FP8 mixed-precision training at that scale; read its recipe before copying the idea.
  6. Fixing the straggler. One slow GPU slows every GPU: Llama 3's authors write that even a single straggler "can slow down thousands of other GPUs." ByteDance's straggler study (May 2025) finds stragglers can come from several causes, not only failed hardware.

Linux and host tuning

Host settings rarely change a kernel's speed. They fix the jitter around it: a GPU that waits on a CPU thread scheduled onto the wrong socket, a NIC interrupt landing on a busy core, a CPU that clocks down between launches. All of these appear as gaps in the Nsight Systems timeline.

SettingWhy it mattersHow to checkRisk and notes
NUMA bindingA process on socket 0 driving a GPU wired to socket 1 crosses the inter-socket link for every copynvidia-smi topo -m for GPU, NIC and CPU affinity; numastat counters numa_miss, other_nodeBind each rank to the CPUs and memory local to its GPU (numactl, or in Slurm --cpus-per-gpu and the Cores= field in gres.conf)
CPU governorpowersave or a slow-ramping governor lets cores clock down between kernel launchesRead scaling_governor per coreThe performance governor requests the highest allowed frequency. With intel_pstate the kernel doc notes it "bypasses the scaling governor layer", so check your driver mode
IRQ affinityNIC interrupts on cores that also run your launch threads add jitter/proc/interrupts; set smp_affinity_list for NIC queuesPin queues to cores local to the NIC. If irqbalance runs, confirm it is not rewriting your choices
Huge pagesFewer TLB misses for large host-memory working sets (pinned buffers, offloaded KV, data loaders)HugePages_Total, HugePages_Free in /proc/meminfoPer the kernel doc, reserve with hugepagesz= and hugepages= at boot or vm.nr_hugepages at runtime. GPU kernels do not use host page tables, so expect effects only on CPU-side work. Measure; some systems prefer transparent huge pages off
GPUDirect RDMANIC reads and writes GPU memory directly, no bounce through host RAMNCCL_DEBUG=INFO shows whether GDR is used; nvidia-peermem loadedThe GPUDirect RDMA docs: devices must share the same upstream PCIe root complex for good results, and it is incompatible with IOMMUs that translate addresses

None of these is a first move. Do them after you have seen the gap in a timeline that they would close.

What to try first

In this order, because each step is cheaper and less risky than the next:

StepChangeExpected effectMeasureWatch for
1Use a current engine and the right attention backend for your GPURemoves avoidable slow pathsKernel names in the profilerSilent fallbacks
2Size the KV cache properly: memory utilisation, max sequences, FP8 KV if quality holdsMore concurrent sequences, fewer preemptionsPreemption counter, p99 latencyQuality on long context
3Tune batching: chunked prefill, max_num_batched_tokensBetter throughput/latency balanceTTFT, inter-token latency, throughput at fixed concurrencyInteractive p99
4Enable prefix caching if prefixes repeatCheaper prefillCache hit rate, TTFTZero gain without repetition
5CUDA graphs and compile on for decodeLower small-batch latencyGaps in Nsight SystemsEager fallbacks, memory
6Weight quantisation (FP8 first on Hopper; W4A16 where memory forces it)Fewer GPUs, faster decodeQuality evals plus latency at your concurrencyAccuracy regressions
7Speculative decoding for low-concurrency latencyLower latency per streamAcceptance rate, latency at low and high loadHurts saturated throughput
8Parallelism layout and NCCLDepends on topologynccl-tests, collective time in NsightTP across slow links
9Host tuning: NUMA, governor, IRQLess jitterTimeline gaps, p99Hidden by earlier steps
10Custom kernelsOnly for a profiled hot operatorNsight Compute rooflineMaintenance cost

Pitfalls we would check first

  • Benchmarking with one prompt length. Real traffic mixes lengths, and the mix changes every conclusion above.
  • Reporting throughput without latency, or the reverse.
  • Comparing against a two-year-old baseline. Paper speedups are against their era's baselines.
  • Changing numerics and performance in the same run, so you cannot tell which broke quality.
  • Reading a vendor figure quoted with sparsity as a dense figure.
  • Tuning one GPU, then adding TP across a slower link and wondering why scaling is poor.

Where Swfte fits

This guide is general engineering advice and does not describe a specific Swfte cluster. If you want to run open-weight models on infrastructure you control, Deploy models and dedicated cloud describe the options, GPU covers compute, and Connect is the model gateway that you can self-deploy. The infrastructure layer page shows how these fit together, and best sovereign cloud providers compares hosting options. For the engine side, see the 2026 serving frameworks comparison and the self-hosted stack guide.

Frequently asked questions

Which optimisation should I try first for a slow LLM endpoint?

Check the basics (attention backend, KV cache size, batching and chunked prefill), then profile. The roofline arithmetic above gives a floor for decode speed; if you are near it, kernel tweaks will not help, and weight compression, batching or more GPUs will.

Does FlashAttention change model outputs?

FlashAttention computes exact attention, not an approximation, so results match the standard implementation up to floating-point differences from reordered arithmetic. Quantisation, FP8 KV cache and some speculative-decoding setups are different: they can change outputs and need evaluation.

Is INT4 always faster than FP8?

No. W4A16 cuts the bytes read per decode step, which helps at small batch. It still computes in 16-bit, so it does little for compute-bound prefill. On Hopper-class GPUs FP8 W8A8 is often the simpler first move. Measure at your own concurrency.

Do CUDA graphs help with long prompts?

Mostly not. They remove CPU launch overhead, which matters when kernels are tiny, as in small-batch decode. Prefill kernels are large and compute-bound, so launch overhead is a small share. Graphs also add memory cost and fall back to eager execution when the batch shape was not captured.

When is a custom Triton or CUTLASS kernel worth it?

When a profiler shows one operator dominating, no engine backend covers it, and you can write a reference test and benchmark across your real shapes. Otherwise the maintenance cost across GPU generations and engine versions outweighs the gain.

All links were read in October 2026. Engine flags and defaults change between releases, so check the version you run.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.