NVIDIA A100 80GB SXM: specs, pricing and what it can run

The previous generation workhorse. No FP8, so modern quantised serving stacks give up a large fraction of their advantage here.

Specifications

SpecNVIDIA A100 80GB SXM
ArchitectureAmpere
Released2020
Memory80 GB HBM2e
Memory bandwidth2,039 GB/s
Dense FP16312 TFLOPS
FP8 tensor coresNo
Board power400 W
InterconnectNVLink 3.0 (600 GB/s)
Segmentdatacenter

What fits in 80 GB

Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.

QuantisationLargest model (weights only)
FP16~36B parameters
FP8 / INT8~73B parameters
4-bit~146B parameters

Rental cost

Indicative on-demand pricing runs roughly $0.80–$2.20 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.

Related

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.