← The journal
Engineering

Why the Cluster Is Rarely the Bottleneck: 10 Mistakes

Ten inference mistakes that waste more GPU money than any hardware shortfall, each with its fix and metric.

Swfte Journal / Engineering

When an inference endpoint is slow or expensive, the first instinct is to buy more GPUs. Sometimes that is right. Far more often the GPUs you already have are idle half the time, waiting on a CPU, stuck behind a long prompt, evicting each other's cache, or running at a batch size that cannot use their compute.

These are the ten mistakes we would check before approving a hardware order, roughly in the order they cost money. The fixes mostly come from published papers and engine documentation, and we link them. Where a number appears, it has a source, hardware and date, or it is labelled as our arithmetic. Companion reads: kernel-level optimisations, building a GPU cluster and serving 400B+ models.

Quick triage

Slow or costly endpoint?
  |
  +-- GPU timeline has gaps?  --> mistake 9 (host and launch overhead)
  +-- Lots of preemptions?    --> mistake 3 (KV cache sizing)
  +-- p99 spikes, p50 fine?   --> mistakes 5 and 6 (prefill stalls, queueing)
  +-- Low tokens/s per GPU?   --> mistakes 1, 2, 4 (workload, batch, prefix reuse)
  +-- Multi-GPU scaling poor? --> mistake 8 (parallelism vs links)
  +-- Cost high, quality fine? -> mistakes 7 and 10 (model size, accounting)

1. Benchmarking a workload you do not have

A single prompt length, a fixed output length and concurrency of one tells you almost nothing. Real traffic has a spread of lengths, bursts, shared prefixes and a mix of interactive and batch callers. Every other conclusion in this list depends on the mix.

Fix: capture a week of real request shapes (input length, output length, arrival times, prefix reuse) and replay them. Report time-to-first-token, inter-token latency at p50 and p99, and throughput at stated concurrency. Metric: your replayed trace's p99, not a vendor chart.

2. Serving at a batch size that cannot use the GPU

Decode is memory-bound. On an H100 SXM, NVIDIA's datasheet lists 3.35 TB/s of bandwidth and a dense BF16 rate of about 989 TFLOPS (half the sparsity figure on the page), a ratio of about 295 FLOPs per byte. A decode step does roughly one FLOP per byte of BF16 weights per sequence in the batch (our arithmetic), so you need on the order of 300 concurrent sequences before the tensor cores are the limit. At batch 8 they are mostly idle, and an extra GPU buys you more idle tensor cores.

Fix: pool traffic onto fewer, busier replicas; use continuous batching (the iteration-level scheduling introduced by Orca, now standard in vLLM, SGLang and TensorRT-LLM); move batch work to a separate offline path, as in our batch inference guide. Metric: achieved tokens per second per GPU against the bandwidth floor, and average running batch size.

3. Sizing for weights and forgetting the KV cache

Weights fit, the demo works, then production traffic arrives with long contexts and the cache fills. When it does, the scheduler preempts sequences and recomputes them, and tail latency jumps.

KV cache per token is 2 x layers x KV heads x head dimension x bytes. With Llama 3 70B's published configuration that is 320 KiB per token in BF16, about 40 GiB for one 128K-token sequence (our arithmetic from the Llama 3 paper). Ten such sequences is 400 GiB.

Fix: size cache for your real context distribution and concurrency. vLLM's optimisation guide lists the levers when preemption shows up: raise gpu_memory_utilization, lower max_num_seqs or max_num_batched_tokens, or raise parallelism so each GPU has more cache. An FP8 KV cache can roughly double token capacity, subject to a quality check. Metric: preemption count, and cache utilisation at peak.

4. Recomputing prefixes you have already seen

Agents, RAG systems and chat apps send the same system prompt, tool definitions or document on every call. Without prefix caching you pay prefill each time. vLLM's automatic prefix caching reuses KV blocks for shared prefixes, and its doc is clear about the limits: it only speeds up prefill, and gives nothing when prefixes differ. SGLang's RadixAttention takes a radix-tree approach for the same goal.

Fix: put stable content first and variable content last in prompts, enable prefix caching, and keep identical tool schemas byte-for-byte stable, since one changed token ahead of the shared part invalidates the match. Metric: prefix-cache hit rate and TTFT for repeated prompts.

5. Letting long prefills stall everyone

A single 30,000-token prompt scheduled with decode traffic can hold every other stream's next token hostage for the duration of its prefill. Sarathi-Serve (March 2024) addresses this with chunked prefills and stall-free scheduling, and reports 2.6 times the serving capacity of vLLM on Mistral-7B on one A100, with larger gains on bigger models, against the vLLM of that time.

Fix: use an engine with chunked prefill. vLLM's V1 engine enables it by default where it can, and its guide says smaller max_num_batched_tokens values (for example 2048) favour inter-token latency while larger ones favour TTFT. Metric: p99 inter-token latency while long prompts are in flight.

6. Judging by averages and ignoring the queue

Average latency hides the experience of your worst users. A GPU at 95 percent utilisation with an unbounded queue is not healthy; it is about to time out. Autoscaling on GPU utilisation is also a poor trigger, because a GPU reports as busy whenever any kernel is running.

Fix: set service objectives on percentiles and on queue time. Scale on queue depth, waiting requests or KV cache pressure. Shed or reroute load before the queue grows past your latency budget. Metric: p99 TTFT, time spent waiting, requests rejected.

7. Using a bigger model or higher precision than the task needs

Many requests are classification, extraction or short rewrites that a much smaller model handles. And on Hopper-class GPUs, FP8 weights halve the bytes read per decode step against BF16. The FP8 formats paper (September 2022) shows FP8 matching 16-bit quality across several architectures, though you should verify on your own tasks. Remember that vLLM lists FP8 W8A8 as supported on Ada, Hopper and AMD GPUs, not A100.

Fix: route by task difficulty (see our router explainer), evaluate quantised variants against your own test set, and keep the large model for the requests that need it. Metric: quality on a fixed eval set, cost per successful task.

8. Choosing parallelism that fights the interconnect

Tensor parallelism needs two all-reduces per layer; stretched across nodes, every layer pays a network round trip twice. NVIDIA's August 2024 analysis on H200 GPUs shows up to 1.5 times higher throughput on Llama 3.1 70B from NVSwitch than from point-to-point links, so even inside a node the fabric matters. Across nodes it matters more.

Fix: keep TP within one NVLink domain, use pipeline parallelism to cross nodes, and add data-parallel replicas for throughput. vLLM's parallelism guide says to set TP to the GPUs per node and PP to the node count. We cover layouts in the 400B+ guide. Metric: nccl-tests bus bandwidth at your group size, and collective time in an Nsight Systems trace.

9. Starving the GPU from the host

Small-batch decode launches hundreds of tiny kernels. If the CPU cannot issue them fast enough, the GPU sits idle between launches. Causes: eager mode instead of CUDA graphs, graph capture falling back to eager for unusual batch shapes (vLLM's CUDA graph doc says unmatched batch descriptors run eagerly), a slow tokeniser or API layer in Python, CPU cores running a power-saving governor, or a process pinned to the wrong NUMA node. The CUDA graphs introduction explains why batching launches into one removes per-kernel CPU cost.

Fix: profile with Nsight Systems and look for gaps between kernels. Enable graphs, set the CPU governor to performance (the kernel docs define it as requesting the highest allowed frequency), and bind each process to the CPUs local to its GPU. Metric: GPU idle fraction between kernels in the timeline.

10. Counting GPU-hours instead of tokens, and having no plan for starts and failures

A GPU at 20 percent average load costs five times as much per token as one at 100 percent (our arithmetic). Teams also forget the costs around the edges: a 400 GB model takes minutes to load, so scale-up is slow; a replica spanning many GPUs dies when one GPU does; and spare capacity for failures has to come from somewhere.

Fix: report cost per million tokens and per successful request, computed from the all-in cost per GPU-hour and measured throughput. Cache weights on local NVMe, keep warm spares, and design for replica loss. The GPU cluster guide has the cost formula and failure table. Metric: cost per million output tokens at observed utilisation, cold-start time, and time to recover from a lost replica.

The ten at a glance

#MistakeSymptomFixMetric
1Wrong benchmarkProduction disappointsReplay real tracesp99 on replay
2Tiny batchesLow tokens/s per GPUPool traffic, continuous batchingRunning batch size
3Cache not sizedPreemptions, p99 spikesSize cache, FP8 KV, parallelismPreemption count
4No prefix reuseHigh TTFT on repeated promptsPrefix caching, stable prompt orderHit rate
5Prefill stalls decodeInter-token spikesChunked prefillp99 ITL
6Averages and queuesTimeouts at peakPercentile SLOs, queue-based scalingQueue time
7Oversized modelHigh cost, fine qualityRoute, quantise, evaluateCost per task
8Bad parallelismPoor multi-GPU scalingTP in NVLink, PP across nodesBus bandwidth
9Host starvationGaps between kernelsCUDA graphs, governor, NUMAGPU idle fraction
10Wrong cost unitSurprise bills, slow recoveryCost per token, warm sparesCost per M tokens

When more hardware is the answer

If the profile shows tensor cores busy, cache comfortably sized, a healthy fabric and queues that still grow, you are out of efficiency to find and need capacity. Then size it with real numbers: our cluster guide covers memory arithmetic and what breaks at scale. For observability of the application layer see LLM observability.

Where Swfte fits

This list is general engineering guidance and does not describe a Swfte cluster. If you want open-weight models running on infrastructure you control, Deploy models, dedicated cloud and GPU are the starting pages; Connect is the model gateway you can self-deploy, and the infrastructure layer shows where each sits. Compare hosting options on best sovereign cloud providers.

Frequently asked questions

Why is my GPU at 100 percent utilisation but throughput is low?

The utilisation counter reports the share of time any kernel was running, not how much of the chip was used. A memory-bound decode kernel at small batch can show 100 percent while using a fraction of the compute. Look at achieved tokens per second against the bandwidth floor, and the average batch size.

How big a batch do I need to use an H100 well in decode?

On the datasheet figures, a dense model needs on the order of 300 concurrent sequences before decode stops being memory-bound (our arithmetic: about 989 TFLOPS dense BF16 divided by 3.35 TB/s). Attention reads and the KV cache shift that figure, so measure at your own context lengths.

Does prefix caching always help?

No. It reduces prefill work when requests share a prefix, and does nothing for generation time or for requests with different prefixes. Check your cache hit rate.

Should I scale on GPU utilisation?

It is a weak signal. Queue depth, waiting requests, time to first token and KV-cache pressure track user experience better, and they react before latency has already degraded.

What should I fix first?

Measure with a realistic traffic replay, then check for gaps in an Nsight Systems timeline and for preemptions. Those two findings usually point to one of the ten above.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.