NVIDIA H100 SXM vs NVIDIA H200 SXM: specs, bandwidth and cost compared

The default training and high-throughput inference part for the 2023–2025 generation, and still the unit most capacity is quoted in. Same Hopper compute as H100 with 76% more memory and 43% more bandwidth. For inference, which is memory-bound, that bandwidth is the whole story.

Specifications

SpecNVIDIA H100 SXMNVIDIA H200 SXM
ArchitectureHopperHopper
Released20222024
Memory80 GB HBM3141 GB HBM3e
Memory bandwidth3,350 GB/s4,800 GB/s
Dense FP16989 TFLOPS989 TFLOPS
FP8 tensor coresYesYes
Board power700 W700 W
InterconnectNVLink 4.0 (900 GB/s)NVLink 4.0 (900 GB/s)
Segmentdatacenterdatacenter

Which one to pick

Inference throughput on large models is bound by memory bandwidth far more than by peak FLOPS. NVIDIA H200 SXM has 4,800 GB/s against 3,350 GB/s, a 1.43× difference, which is the figure that most closely tracks tokens per second at batch size 1. Training and large-batch serving lean more on FP16 throughput, where the gap is 1.00×.

Rental cost

Indicative on-demand pricing runs roughly $1.90–$4.50 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.