NVIDIA H200 SXM: specs, pricing and what it can run
Same Hopper compute as H100 with 76% more memory and 43% more bandwidth. For inference, which is memory-bound, that bandwidth is the whole story.
Specifications
| Spec | NVIDIA H200 SXM |
|---|---|
| Architecture | Hopper |
| Released | 2024 |
| Memory | 141 GB HBM3e |
| Memory bandwidth | 4,800 GB/s |
| Dense FP16 | 989 TFLOPS |
| FP8 tensor cores | Yes |
| Board power | 700 W |
| Interconnect | NVLink 4.0 (900 GB/s) |
| Segment | datacenter |
What fits in 141 GB
Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.
| Quantisation | Largest model (weights only) |
|---|---|
| FP16 | ~64B parameters |
| FP8 / INT8 | ~128B parameters |
| 4-bit | ~257B parameters |
Rental cost
Indicative on-demand pricing runs roughly $2.50–$6.00 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.