NVIDIA A100 80GB SXM vs NVIDIA H100 SXM: specs, bandwidth and cost compared

The previous generation workhorse. No FP8, so modern quantised serving stacks give up a large fraction of their advantage here. The default training and high-throughput inference part for the 2023–2025 generation, and still the unit most capacity is quoted in.

Specifications

SpecNVIDIA A100 80GB SXMNVIDIA H100 SXM
ArchitectureAmpereHopper
Released20202022
Memory80 GB HBM2e80 GB HBM3
Memory bandwidth2,039 GB/s3,350 GB/s
Dense FP16312 TFLOPS989 TFLOPS
FP8 tensor coresNoYes
Board power400 W700 W
InterconnectNVLink 3.0 (600 GB/s)NVLink 4.0 (900 GB/s)
Segmentdatacenterdatacenter

Which one to pick

Inference throughput on large models is bound by memory bandwidth far more than by peak FLOPS. NVIDIA H100 SXM has 3,350 GB/s against 2,039 GB/s, a 1.64× difference, which is the figure that most closely tracks tokens per second at batch size 1. Training and large-batch serving lean more on FP16 throughput, where the gap is 3.17×.

Rental cost

Indicative on-demand pricing runs roughly $0.80–$2.20 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.