NVIDIA GeForce RTX 5090 vs NVIDIA H100 SXM: specs, bandwidth and cost compared

The strongest consumer option for local inference: 32GB of GDDR7 at 1.79 TB/s beats every previous consumer part. Licence terms restrict datacenter deployment. The default training and high-throughput inference part for the 2023–2025 generation, and still the unit most capacity is quoted in.

Specifications

SpecNVIDIA GeForce RTX 5090NVIDIA H100 SXM
ArchitectureBlackwellHopper
Released20252022
Memory32 GB GDDR780 GB HBM3
Memory bandwidth1,792 GB/s3,350 GB/s
Dense FP16419 TFLOPS989 TFLOPS
FP8 tensor coresYesYes
Board power575 W700 W
InterconnectPCIe Gen5 (no NVLink)NVLink 4.0 (900 GB/s)
Segmentconsumerdatacenter

Which one to pick

Inference throughput on large models is bound by memory bandwidth far more than by peak FLOPS. NVIDIA H100 SXM has 3,350 GB/s against 1,792 GB/s, a 1.87× difference, which is the figure that most closely tracks tokens per second at batch size 1. Training and large-batch serving lean more on FP16 throughput, where the gap is 2.36×.

Rental cost

Indicative on-demand pricing runs roughly $0.40–$1.20 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check retail + live cloud price feed before budgeting.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.