NVIDIA B200 vs NVIDIA H200 SXM: specs, bandwidth and cost compared
Blackwell generation. Roughly 2.3× H100 dense FP16 and 2.4× the memory bandwidth, with native FP4 for inference. Same Hopper compute as H100 with 76% more memory and 43% more bandwidth. For inference, which is memory-bound, that bandwidth is the whole story.
Specifications
| Spec | NVIDIA B200 | NVIDIA H200 SXM |
|---|---|---|
| Architecture | Blackwell | Hopper |
| Released | 2025 | 2024 |
| Memory | 192 GB HBM3e | 141 GB HBM3e |
| Memory bandwidth | 8,000 GB/s | 4,800 GB/s |
| Dense FP16 | 2,250 TFLOPS | 989 TFLOPS |
| FP8 tensor cores | Yes | Yes |
| Board power | 1000 W | 700 W |
| Interconnect | NVLink 5.0 (1.8 TB/s) | NVLink 4.0 (900 GB/s) |
| Segment | datacenter | datacenter |
Which one to pick
Inference throughput on large models is bound by memory bandwidth far more than by peak FLOPS. NVIDIA B200 has 8,000 GB/s against 4,800 GB/s, a 1.67× difference, which is the figure that most closely tracks tokens per second at batch size 1. Training and large-batch serving lean more on FP16 throughput, where the gap is 2.28×.
Rental cost
Indicative on-demand pricing runs roughly $4.00–$11.00 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.