NVIDIA GeForce RTX 5090 vs NVIDIA H100 SXM: specs, bandwidth and cost compared
The strongest consumer option for local inference: 32GB of GDDR7 at 1.79 TB/s beats every previous consumer part. Licence terms restrict datacenter deployment. The default training and high-throughput inference part for the 2023–2025 generation, and still the unit most capacity is quoted in.
Specifications
| Spec | NVIDIA GeForce RTX 5090 | NVIDIA H100 SXM |
|---|---|---|
| Architecture | Blackwell | Hopper |
| Released | 2025 | 2022 |
| Memory | 32 GB GDDR7 | 80 GB HBM3 |
| Memory bandwidth | 1,792 GB/s | 3,350 GB/s |
| Dense FP16 | 419 TFLOPS | 989 TFLOPS |
| FP8 tensor cores | Yes | Yes |
| Board power | 575 W | 700 W |
| Interconnect | PCIe Gen5 (no NVLink) | NVLink 4.0 (900 GB/s) |
| Segment | consumer | datacenter |
Which one to pick
Inference throughput on large models is bound by memory bandwidth far more than by peak FLOPS. NVIDIA H100 SXM has 3,350 GB/s against 1,792 GB/s, a 1.87× difference, which is the figure that most closely tracks tokens per second at batch size 1. Training and large-batch serving lean more on FP16 throughput, where the gap is 2.36×.
Rental cost
Indicative on-demand pricing runs roughly $0.40–$1.20 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check retail + live cloud price feed before budgeting.