NVIDIA B200: specs, pricing and what it can run
Blackwell generation. Roughly 2.3× H100 dense FP16 and 2.4× the memory bandwidth, with native FP4 for inference.
Specifications
| Spec | NVIDIA B200 |
|---|---|
| Architecture | Blackwell |
| Released | 2025 |
| Memory | 192 GB HBM3e |
| Memory bandwidth | 8,000 GB/s |
| Dense FP16 | 2,250 TFLOPS |
| FP8 tensor cores | Yes |
| Board power | 1000 W |
| Interconnect | NVLink 5.0 (1.8 TB/s) |
| Segment | datacenter |
What fits in 192 GB
Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.
| Quantisation | Largest model (weights only) |
|---|---|
| FP16 | ~87B parameters |
| FP8 / INT8 | ~175B parameters |
| 4-bit | ~350B parameters |
Rental cost
Indicative on-demand pricing runs roughly $4.00–$11.00 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.