GPU comparison for LLM inference and training

Token generation on a large model is bound by memory bandwidth long before it is bound by peak FLOPS. The two columns worth reading first are memory and bandwidth; everything else is secondary for inference.

GPUMemoryBandwidthFP16Max 4-bit model$/GPU-hr
NVIDIA B200192 GB HBM3e8,000 GB/s2,250 TFLOPS~350B$4.00–$11.00
AMD Instinct MI300X192 GB HBM35,300 GB/s1,307 TFLOPS~350B$1.80–$4.50
NVIDIA H200 SXM141 GB HBM3e4,800 GB/s989 TFLOPS~257B$2.50–$6.00
NVIDIA H100 SXM80 GB HBM33,350 GB/s989 TFLOPS~146B$1.90–$4.50
NVIDIA A100 80GB SXM80 GB HBM2e2,039 GB/s312 TFLOPS~146B$0.80–$2.20
NVIDIA GeForce RTX 509032 GB GDDR71,792 GB/s419 TFLOPS~58B$0.40–$1.20
NVIDIA GeForce RTX 409024 GB GDDR6X1,008 GB/s330 TFLOPS~43B$0.30–$0.90
NVIDIA L40S48 GB GDDR6864 GB/s362 TFLOPS~87B$0.70–$1.80

Rental figures are indicative on-demand ranges across major clouds, not quotes. Spot and reserved capacity price well below the low end.

Head-to-head

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.