GPU comparison for LLM inference and training
Token generation on a large model is bound by memory bandwidth long before it is bound by peak FLOPS. The two columns worth reading first are memory and bandwidth; everything else is secondary for inference.
| GPU | Memory | Bandwidth | FP16 | Max 4-bit model | $/GPU-hr |
|---|---|---|---|---|---|
| NVIDIA B200 | 192 GB HBM3e | 8,000 GB/s | 2,250 TFLOPS | ~350B | $4.00–$11.00 |
| AMD Instinct MI300X | 192 GB HBM3 | 5,300 GB/s | 1,307 TFLOPS | ~350B | $1.80–$4.50 |
| NVIDIA H200 SXM | 141 GB HBM3e | 4,800 GB/s | 989 TFLOPS | ~257B | $2.50–$6.00 |
| NVIDIA H100 SXM | 80 GB HBM3 | 3,350 GB/s | 989 TFLOPS | ~146B | $1.90–$4.50 |
| NVIDIA A100 80GB SXM | 80 GB HBM2e | 2,039 GB/s | 312 TFLOPS | ~146B | $0.80–$2.20 |
| NVIDIA GeForce RTX 5090 | 32 GB GDDR7 | 1,792 GB/s | 419 TFLOPS | ~58B | $0.40–$1.20 |
| NVIDIA GeForce RTX 4090 | 24 GB GDDR6X | 1,008 GB/s | 330 TFLOPS | ~43B | $0.30–$0.90 |
| NVIDIA L40S | 48 GB GDDR6 | 864 GB/s | 362 TFLOPS | ~87B | $0.70–$1.80 |
Rental figures are indicative on-demand ranges across major clouds, not quotes. Spot and reserved capacity price well below the low end.