← The journal
Engineering

Building a GPU Cluster for Large Models, and What Breaks

Memory maths, interconnect, storage, failures, schedulers and power for large-model GPU clusters, plus a checklist.

Swfte Journal / Engineering

A GPU cluster for large models is a memory problem first, a network problem second and an operations problem for the rest of its life. The purchase order is the easy part. The hard part is that thousands of components have to work at the same moment, and a synchronous training job stops when any one of them does not.

This guide covers sizing, interconnect, storage, the failures you should expect, scheduling, power, security and the build-versus-rent decision, and ends with a pre-flight checklist and a failure-mode table. Numbers come from cited primary sources with their hardware and date. Where we do arithmetic ourselves, we say so. We do not quote prices or delivery times, because we have no source we could verify for either.

Companion reads: kernel-level optimisations, serving 400B+ models and the ten inference mistakes that cost the most.

Decide what the cluster is for

Training and serving want different machines. A rough split, which is a judgement and not a rule:

  • Pre-training or large fine-tuning: synchronous, network-hungry, checkpoint-heavy. Scale-out fabric quality decides your utilisation.
  • Serving: latency-sensitive, bursty, lower inter-node traffic unless the model spans nodes. Memory capacity and HBM bandwidth decide cost per token.
  • Fine-tuning small and mid-size models: often fits in one or two nodes, and does not need a fabric project.

If you are not pre-training a frontier-scale model, you probably do not need a thousand-GPU fabric. Many teams that think they need a cluster need one node with enough memory and a good serving engine.

Sizing: the memory arithmetic

Four things consume GPU memory: weights, KV cache, activations and, in training, optimiser state and gradients.

Weights are parameters times bytes per parameter. This is arithmetic, with parameter counts from the model papers:

ModelParametersBF16 (2 B)FP8 / INT8 (1 B)INT4 (about 0.5 B plus scales)
Llama 3 70B70B140 GB70 GBabout 35 GB
Llama 3.1 405B405B810 GB405 GBabout 203 GB
Mixtral 8x7B47B total, 13B active94 GB47 GBabout 24 GB
DeepSeek-V3671B total, 37B active1.34 TB671 GBabout 336 GB

MoE models need memory for all experts but spend compute on the active ones, so memory sizing follows total parameters. Sources: Llama 3, Mixtral, DeepSeek-V3.

What a node holds. The DGX H100 user guide lists 8 H100 GPUs and 640 GB total GPU memory. NVIDIA's H200 page lists 141 GB per GPU, so an 8-GPU H200 node holds 1,128 GB (our arithmetic). Putting those together:

  • Llama 3 70B in BF16 (140 GB) fits on two 80 GB GPUs with a little room, and comfortably on four.
  • Llama 3.1 405B in FP8 (405 GB) fits on one 8x H100 node with about 235 GB left for KV cache, activations and engine overhead. In BF16 (810 GB) it does not fit on one H100 node and needs two or an H200 node.
  • DeepSeek-V3 in FP8 (671 GB) does not fit in 640 GB. It needs two H100 nodes or one H200 node (1,128 GB), and real deployments use far more GPUs for throughput. DeepSeek's own decode deployment uses 320 GPUs as the minimum unit, per its technical report.

KV cache is 2 x layers x KV heads x head dimension x bytes, per token, per sequence. Using Llama 3 hyperparameters, that is 320 KiB per token for 70B and 504 KiB for 405B in BF16, which is about 40 GiB and 63 GiB for one 128K-token sequence. We derive this in the kernel guide. Sizing the cache for your real context-length distribution and concurrency is the step most capacity plans skip.

Training state. The ZeRO paper counts mixed-precision Adam at 16 bytes per parameter: 2 for FP16 weights, 2 for gradients and 12 for FP32 weights, momentum and variance. Our arithmetic:

ModelState at 16 B/paramMinimum 80 GB GPUs for state alone
70B1.12 TB14
405B6.48 TB81
671B10.7 TB135

That is before activations and is the floor, reached only if state is sharded perfectly (ZeRO or FSDP). Training recipes differ, and some keep states in lower precision, so treat this as a planning default. Llama 3 405B used 4D parallelism (tensor, context, pipeline, data) so that "each GPU's model parameters, optimizer states, gradients, and activations fit in its HBM."

Time to train. The Llama 3 paper gives 3.8 x 10^25 FLOPs for the 405B pre-training run and 400 TFLOPs per GPU at 16,384 GPUs in Table 4. Dividing, the arithmetic is about 5.8 million seconds, or 67 days at that rate. That is a sanity check of scale, not a claim about how long their run took. Use the same method with your own token budget and a realistic utilisation. The paper's BF16 model FLOPs utilisation was 38 to 43 percent.

There are two networks, and they are different jobs.

Inside the node. NVLink and NVSwitch. NVIDIA lists 900 GB/s of NVLink bandwidth per H100, and the DGX H100 guide lists 900 GB/s GPU-to-GPU bandwidth inside the node. This is where tensor parallelism and expert-parallel traffic should live. With Blackwell's GB200 NVL72, the NVLink domain grows to 72 GPUs in one liquid-cooled rack, which moves the boundary between "fast" and "slow" parallelism and changes what you can place on NVLink.

Between nodes. InfiniBand or RDMA over Converged Ethernet (RoCE), typically one 400 Gb/s NIC per GPU. Meta's Llama 3 paper is the best public comparison: the 405B model trained on a RoCE fabric built from Arista and Minipack2 switches, the smaller models on NVIDIA Quantum-2 InfiniBand, "both RoCE and InfiniBand clusters leverage 400 Gbps interconnects" and were tuned to give equivalent performance for these workloads.

InfiniBandRoCE (Ethernet)
StrengthsMature RDMA fabric, well-trodden by NVIDIA's reference designsEthernet operations skills and tooling, multi-vendor switches, can share fabric tooling with the rest of the site
CostsSeparate fabric and skills, single-vendor ecosystemYou own congestion control, load balancing and tuning
EvidenceNVIDIA SuperPOD reference designsMeta's published RoCE experience

Meta's account of making RoCE work is instructive. Plain ECMP gave poor load balancing because AI traffic has low flow entropy. Path pinning degraded performance by more than 30 percent when job placement was fragmented. Enhanced ECMP hashing on the RoCE queue-pair field gave up to 40 percent better AllReduce. They disabled DCQCN at 400G and moved to receiver-driven admission built into the collective library. The lesson they draw: success needed "deep coordination between the collective communication library and the network." RoCE is viable, and the work lands on you.

Rail-optimised topology. In NVIDIA's DGX SuperPOD H100 reference architecture, each group of 32 nodes is "rail-aligned": the same-numbered NIC on every node connects to the same leaf switch, so traffic between GPU k on one node and GPU k on another is one hop.

node A: GPU0 GPU1 ... GPU7        node B: GPU0 GPU1 ... GPU7
          |    |        |                   |    |        |
         NIC0 NIC1     NIC7                NIC0 NIC1     NIC7
          |    |        |                   |    |        |
       [leaf0][leaf1] ..[leaf7]  <- one leaf switch per rail
              \    |    /
               [ spine layer ]   <- only for traffic that changes rail

Collectives like AllReduce largely stay on a rail, so the spine carries less. Check that your scheduler places jobs to keep ranks on matching rails, or you pay for the fabric and do not get its benefit.

Oversubscription. Meta's 24K-GPU RoCE cluster is a three-layer Clos with full bisection bandwidth inside a 3,072-GPU pod, and a 1:7 oversubscription at the aggregation layer, so its parallelism and scheduler are "optimized to be aware of network topology, aiming to minimize network communication across pods." A cheaper fabric is a legitimate choice if your software knows where the thin parts are.

What to test on day one: nccl-tests (all_reduce_perf, all_to_all_perf) across 2, 8, 64 nodes and your largest job size, comparing bus bandwidth to the expected line rate, and per-link error and retransmit counters. Fix the fabric before debugging model code.

Storage and data loading

Training reads data continuously and writes checkpoints in bursts. Meta's Llama 3 storage was Tectonic: 240 PB across 7,500 SSD servers, 2 TB/s sustained and 7 TB/s peak, and the paper calls "highly bursty checkpoint writes that saturate the storage fabric for short durations" a major challenge. Most teams are three orders of magnitude smaller, but the shape of the problem is the same.

  • Data loading. Many small files, random access over a network filesystem, and CPU-bound decoding or tokenisation starve GPUs quietly. Pre-tokenise into large sharded files, stream sequentially, and measure input wait in the profiler. DGX H100 ships with eight 3.84 TB NVMe drives as a local cache, which is a hint about where its designers expect hot data to live.
  • Checkpointing. PyTorch's Distributed Checkpoint writes from many ranks in parallel, reshards on load when topology changes and offers async_save so training continues while the write finishes. Decide checkpoint frequency from the cost of lost work against the cost of a pause. Test restoring onto a different node count, because after a failure you may not get the same machines back.
  • Direct paths. GPUDirect Storage removes the CPU bounce buffer between storage and GPU memory for supported filesystems. Worth it for load and restore, not a first priority.
  • Serving. Cold start is a storage problem: a 400 GB checkpoint pulled over a slow link makes every scale-up take minutes. Cache weights on local NVMe or a nearby object store.

What goes wrong at scale

Meta's Llama 3 paper is the most useful public failure log. During a 54-day snapshot of 405B pre-training on up to 16K H100 GPUs, there were 466 job interruptions: 47 planned (maintenance, firmware, configuration) and 419 unexpected. About 78 percent of the unexpected ones were attributed to confirmed or suspected hardware issues, and GPU issues were the largest category at 58.7 percent. The largest rows in Table 5 were faulty GPUs (148), GPU HBM3 memory (72), software bugs (54) and network switches or cables (35). Only three needed significant manual intervention, the rest were handled by automation, and effective training time stayed above 90 percent.

Read that carefully. Even a very well-run cluster failed more than seven times a day on average, and hardware was most of it. What kept it productive was automation: fast detection, fast restart, fast diagnosis. If your plan assumes failures are rare, the plan is wrong.

Stragglers and hangs

A synchronous job runs at the speed of its slowest rank. The Llama 3 authors write that "even a single straggler can slow down thousands of other GPUs, often appearing as functioning but slow communications." ByteDance's analysis of five months of cluster traces (May 2025) finds stragglers are not always hardware failures and can come from several causes. Typical sources: a GPU throttling on heat or power, a degraded link, a noisy neighbour, uneven pipeline stages, a slow dataloader worker.

Hangs are worse than slowness because nothing reports an error. Llama 3's paper notes that a failing NVLink "often manifest[s] as stalled load/store operations within CUDA kernels without returning a clear error code." The countermeasures are the same everywhere:

  • Run PyTorch's NCCL watchdog and flight recorder. TORCH_NCCL_TRACE_BUFFER_SIZE sets the flight-recorder ring-buffer size, TORCH_NCCL_DUMP_ON_TIMEOUT writes debug data when the watchdog fires, TORCH_NCCL_ASYNC_ERROR_HANDLING controls teardown, and TORCH_NCCL_HEARTBEAT_TIMEOUT_SEC sets when the monitor aborts a stuck process. Meta says it relies on the flight recorder to "diagnose hangs and performance issues quickly at scale."
  • Set timeouts on purpose. Too short and slow checkpoint or startup phases kill healthy jobs; too long and a dead rank blocks the whole allocation for hours.
  • Log per-rank step time so you can find the slow rank by looking, not guessing.
  • Quarantine a node the moment it is implicated, and check it before returning it to the pool.

NCCL and fabric problems

Symptoms: a job that scales fine to 8 nodes and collapses at 64, intermittent timeouts, bandwidth far below line rate. Causes we would check in order: wrong interface or HCA selection (NCCL_SOCKET_IFNAME, NCCL_IB_HCA), GPUDirect RDMA silently not in use, PCIe ACS or IOMMU blocking peer-to-peer, a bad cable or optic showing as retransmits, ECMP imbalance, MTU or PFC and ECN mismatches on RoCE. NCCL_DEBUG=INFO for one run shows which transport NCCL chose, and the NCCL troubleshooting guide lists PCI ACS, IOMMU, shared-memory and RoCE diagnostics.

ECC errors and Xid messages

GPU faults surface as Xid messages in the kernel log. NVIDIA's Xid catalogue is the reference. A few that matter, with the catalogue's own descriptions:

XidMeaningDocumented response
13Graphics engine exception, usually an application out-of-bounds errorRestart the application; debug with Compute Sanitizer
31GPU memory page fault from an illegal address accessRestart; debug, or contact support if inconclusive
48Double-bit ECC error, uncorrectableGPU reset or node reboot; check thresholds for RMA
63Memory row-remapping eventInformational
64Row-remapping failureReset the GPU; contact support if it persists
74NVLink errorFollow the NVLink workflow; may be hardware or the remote device
79GPU has fallen off the busRestart the system; review PCI event logs
94 / 95Contained / uncontained memory errorRestart the application / reset the GPU before restarting applications

On Ampere and newer, row remapping replaces degrading memory cells with spare rows, and the remap needs a GPU reset to take effect. The operational consequence: a node showing pending remaps or repeated Xid 48 or 94 should drain and reset, not keep running training jobs. DCGM diagnostics (dcgmi diag -r 1 to -r 4) range from a quick software, PCIe and memory check to longer hardware stress runs, and NVIDIA recommends the shorter levels as pre-job health checks. Run -r 3 or higher on every new node before accepting it, and after repairs.

Also plan for silent data corruption. Table 5 lists six SDC events among Llama 3's 419 unexpected interruptions. Loss spikes with no software cause are a hint; keep known-good reference runs and compare.

Thermal and power limits

H100 SXM modules are rated for up to 700 W (configurable), per NVIDIA's product page. The DGX H100 guide lists a 10.2 kW maximum for the system and a 5 to 30 degrees Celsius operating range. Our arithmetic: four such systems in a rack is about 41 kW before network gear, well beyond a conventional enterprise rack budget.

Heat and power also show up in throughput. Llama 3's authors saw a 1 to 2 percent diurnal throughput variation because mid-day temperatures affect GPU clock scaling. They also report that when tens of thousands of GPUs change power draw at once, for example while waiting on a checkpoint, instantaneous swings across the data centre can be "on the order of tens of megawatts, stretching the limits of the power grid." At smaller scale the same effect appears as breaker trips and UPS alarms on synchronised start-up. Stagger job starts, and check what your power provider tolerates.

Scheduler: Slurm, Kubernetes or both

SlurmKubernetes
GPU modelGRES: gres.conf, AutoDetect=nvml, --gres=gpu:N, per-GPU CPU and memory binding, cgroup device isolation (ConstrainDevices), MIG since 21.08Vendor device plugins expose nvidia.com/gpu; GPUs go in limits only and are whole units unless you add MIG or time-slicing
StrengthBuilt for tightly coupled batch jobs, topology-aware placement, mature HPC accountingServices, autoscaling, rolling updates, a huge ecosystem for serving
GapNot a service platformGang scheduling and quota need extras such as Kueue, which adds job queueing, quota sharing, preemption and all-or-nothing admission
Usually fitsLarge training and research clustersInference fleets, platform teams, mixed workloads

A common, defensible split is Slurm (or Kubernetes with Kueue) for training and Kubernetes for serving, on separate node pools so a training burst cannot starve latency-sensitive replicas. Whichever you choose, topology awareness is the thing to test: a job placed across the wrong rails or sockets will underperform and look like a software bug.

Observability

You need four signals, correlated by node and job:

  1. GPU health: DCGM metrics (utilisation, SM clocks, temperature, power, ECC counters, NVLink errors) plus Xid events from the kernel log.
  2. Network: per-port throughput, errors, retransmits and congestion marks on every NIC and switch port.
  3. Job: per-rank step time, MFU or tokens per second, data-loader wait, checkpoint duration, collective timings from the flight recorder.
  4. Facility: inlet temperature, rack power, PDU alarms.

Alert on trends as well as faults: a GPU whose clocks sag at the same time every afternoon, a link whose error counter climbs slowly. Utilisation numbers are cheap to misread. A GPU shows 100 percent "utilisation" if any kernel is running, so track achieved FLOPs or tokens per second instead.

Cost

Do not start from the GPU price. Start from tokens or training runs.

serving cost per 1M output tokens
  = (all-in cost per GPU-hour x GPUs per replica)
    / (output tokens per second per replica x 3600) x 1,000,000

all-in cost per GPU-hour
  = (hardware amortisation + power + cooling + space + network + storage
     + staff + spares + software) / (hours x average utilisation)

Utilisation is the lever. An H100 node at 30 percent average load costs three times as much per token as the same node at 90 percent. Four costs get left out most often: idle time between training jobs, failed-run waste (restart time plus lost steps), spare capacity for failures, and the engineers who run it. For training, count the effective training time, not elapsed time. Llama 3 reports above 90 percent, which is a result of deliberate engineering, not a default.

Procurement and lead times

We have no verified source for current delivery times, and they move with generation and demand, so we will not quote any. What we can say is what to ask for in writing:

  • Delivery dates per component, not per order: GPUs, optics, cables, switches, power distribution and cooling equipment have different lead times, and the slowest sets your date.
  • The power and cooling spec per rack, and who provides what at the rack boundary.
  • Firmware and driver support commitments, RMA turnaround and the spares pool.
  • A burn-in and acceptance procedure, including the DCGM and nccl-tests runs above, with agreed pass criteria.

The GPU procurement strategy post covers the vendor side.

Multi-tenant isolation

Whole-GPU allocation is the strongest isolation you get cheaply: one tenant per GPU, enforced by the scheduler (Slurm cgroup device constraints, Kubernetes device-plugin limits). Beyond that:

  • MIG partitions a GPU into up to seven instances; the H100 datasheet lists "up to 7 MIGs @ 10GB each". It gives hardware-level partitioning of memory and compute, which suits small inference workloads and not large-model training.
  • Shared NVLink domains mean tenants on one NVSwitch fabric share more than you might assume. Decide whether your threat model allows that, and for sensitive work keep tenants on separate nodes.
  • Network isolation needs fabric-level partitioning, not just VLANs on the management network. Ask your fabric vendor what the isolation boundary is and how it is verified.
  • Confidential computing adds hardware-based memory protection for GPU workloads; our explainer covers what it does and does not protect.

Security and air-gapped constraints

An air-gapped cluster changes day-two operations more than day one. You need an internal mirror for OS packages, container images, drivers, firmware and model weights, a controlled path for moving them in, and a way to verify what you moved. Our air-gapped deployment checklist, DMZ architecture guide and self-hosted LLM stack guide cover the application side. For the cluster itself: management networks (BMC, switch consoles) are the usual weak point, firmware updates need a tested rollback, and every extra tool in the pipeline (monitoring, tracing, licence checks) is a potential outbound call to find and block.

EU data-centre and power constraints

If you build or colocate in the EU, power and reporting constrain you before hardware does.

  • Reporting. Under the Energy Efficiency Directive (EU) 2023/1791 and Commission Delegated Regulation (EU) 2024/1364, operators of data centres with installed IT power demand of at least 500 kW report energy, water and other sustainability indicators to a European database. A GPU cluster crosses that line quickly: at the 10.2 kW system figure above, about 50 DGX H100-class systems approach it (our arithmetic). Check whether you or your colocation provider is the reporting party.
  • Grid access. Connection capacity is the scarce resource in some regions. Ireland's regulator CRU published a decision on connection policy for large energy users, mostly data centres, on 12 December 2025 (CRU/2025236, as summarised by William Fry; we read the summary, not the decision itself). Check the rules for your chosen country and site before signing anything.
  • Cooling. Air cooling stops being practical at some rack density, and liquid cooling changes the building, not just the rack. NVIDIA describes GB200 NVL72 as liquid-cooled.
  • Residency. Where weights, data and logs live is a separate question from where GPUs sit; see what in-region must mean and deploying open-source LLMs in the EU.

Build, rent or sovereign-hosted

Build and ownRent (hyperscaler or GPU cloud)Sovereign-hosted (dedicated, region-pinned, operated for you)
ControlHighestLowest to mediumHigh over location and tenancy; operations shared
Time to first jobLongest: site, power, procurement, burn-inShortestMedium
StaffingYour team runs hardware, fabric, firmwareProvider runs hardware; you run softwareProvider or partner runs hardware; agree who owns what
Cost shapeCapital plus fixed operating cost; best at sustained high utilisationPay for use; strongest for spiky or exploratory workContracted capacity; check what is metered
Main riskUnder-utilisation and failures you must absorbCapacity availability, jurisdiction of the operatorContract terms, exit, who has administrative access

Rules of thumb, which are judgement and not measurements: rent while the workload is uncertain; build when utilisation will stay high for years and you can staff it; choose sovereign-hosted when jurisdiction or operator access is the binding constraint and you do not want to run a data centre. For a comparison of providers, see best sovereign cloud providers.

Pre-flight checklist

Before the first real job:

  • Memory plan written: weights, KV cache at real context and concurrency, optimiser state, activations, headroom.
  • Parallelism layout chosen so that tensor and expert traffic stays inside the NVLink domain.
  • Fabric tested with nccl-tests at 2, 8 and full node counts; bus bandwidth recorded against expected line rate.
  • GPUDirect RDMA confirmed in use (NCCL_DEBUG=INFO), including inside containers.
  • BIOS, driver, CUDA, NCCL, firmware and kernel versions pinned and identical across nodes.
  • DCGM diagnostics run on every node at level 3 or higher; failures replaced before the job starts.
  • NUMA, CPU governor and IRQ affinity checked per node.
  • Scheduler placement verified to be topology-aware (rails, sockets, NVLink domain).
  • Flight recorder, watchdog timeouts and per-rank logging enabled.
  • Checkpoint written and restored on a different node count; time recorded.
  • Automatic drain and quarantine for nodes with Xid or ECC alarms.
  • Power and thermal limits confirmed with the facility, including start-up behaviour.
  • Spares, RMA path and on-call named.
  • Isolation boundaries documented for each tenant type.
  • Egress rules and mirrors tested if air-gapped.
  • Cost model set up to report cost per token or per run, not per GPU.

Failure-mode table

SymptomLikely causeFirst checkFix
Job hangs, no errorDead or stalled rank, NVLink stall, fabric faultFlight-recorder dump, last collective per rankQuarantine node, restart from checkpoint
Steps slow, one rank behindStraggler: thermal or power throttle, degraded linkPer-rank step time, GPU clocks, link countersDrain node, investigate, repair
Scaling collapses past N nodesFabric oversubscription, ECMP imbalance, wrong rail placementnccl-tests at N, switch countersTopology-aware placement, load-balancing fixes
NCCL timeout at startupInterface or HCA selection, firewall, GDR offNCCL_DEBUG=INFOCorrect NCCL_SOCKET_IFNAME and NCCL_IB_HCA, open ports
Xid 48 / repeated 94Uncorrectable memory errorsKernel log, nvidia-smi remap statusDrain, reset GPU, RMA if it persists
Xid 79GPU fell off the busPCI event logs, power, seatingReboot, then replace if repeated
Xid 74NVLink errorNVLink counters on both endsFollow NVLink workflow, replace board or baseboard
Loss spike with no code causeSilent data corruption or a bad batchRe-run on a different node; compareQuarantine suspect node; keep reference runs
Throughput dips every afternoonThermal throttlingInlet temperature, SM clocksCooling fix, power-cap review
Checkpoint stalls trainingStorage fabric saturated by bursty writesWrite throughput, GPU idle during saveAsync checkpointing, stagger writes, faster tier
Serving cold start takes minutesWeights pulled over slow pathPull time vs weight sizeLocal NVMe cache, prefetch
Tenant interferenceShared GPU, link or filesystemPer-tenant metricsWhole-GPU allocation, quotas, separate pools

Where Swfte fits

This guide is general and does not describe a Swfte cluster. If you would rather not run one, Deploy models covers running open-weight models, dedicated cloud and GPU describe dedicated capacity, Connect is the model gateway you can self-deploy, and the infrastructure layer explains how they fit. Swfte does not claim to operate its own data centres, so the jurisdiction you get is that of the infrastructure underneath. Compare options on best sovereign cloud providers.

Frequently asked questions

How many GPUs do I need to serve a 400B-class model?

By memory: at FP8, a 405B dense model needs about 405 GB for weights, so one 8x H100 node (640 GB) is the minimum, with roughly 235 GB for cache and overhead. At BF16 you need about 810 GB, so two H100 nodes or one 8x H200 node. Throughput targets usually push beyond the memory minimum. See the 400B+ serving guide.

Is InfiniBand better than RoCE for AI training?

Neither is categorically better. Meta's Llama 3 paper trained its 405B model on RoCE and smaller models on InfiniBand, with both at 400 Gb/s and tuned for equivalent performance. InfiniBand comes with mature reference designs; RoCE lets you use Ethernet skills and switches but puts congestion control and load balancing on you.

How often do GPU clusters fail?

At 16K H100 GPUs, Meta reported 466 job interruptions over a 54-day snapshot, 419 of them unexpected, with GPUs the largest cause. A smaller cluster fails less often in absolute terms but not rarely. Design for automatic detection, restart and node quarantine from the start.

Should I use Slurm or Kubernetes for GPUs?

Slurm suits large, tightly coupled training jobs with topology-aware placement. Kubernetes suits serving, autoscaling and platform workflows, and needs Kueue or similar for queueing and gang scheduling. Many sites run both on separate pools.

Is it cheaper to build or rent?

It depends on utilisation and staffing. Owning wins when GPUs stay busy for years and you can run the hardware; renting wins for uncertain or spiky workloads. Compute cost per token or per training run at your expected utilisation, with failure waste and staff included, before deciding.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.