NVIDIA L40S: specs, pricing and what it can run
Cost-efficient inference and fine-tuning part. GDDR6 bandwidth is a quarter of H100, so it suits smaller models and batch work rather than latency-critical serving.
Specifications
| Spec | NVIDIA L40S |
|---|---|
| Architecture | Ada Lovelace |
| Released | 2023 |
| Memory | 48 GB GDDR6 |
| Memory bandwidth | 864 GB/s |
| Dense FP16 | 362 TFLOPS |
| FP8 tensor cores | Yes |
| Board power | 350 W |
| Interconnect | PCIe Gen4 (no NVLink) |
| Segment | datacenter |
What fits in 48 GB
Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.
| Quantisation | Largest model (weights only) |
|---|---|
| FP16 | ~21B parameters |
| FP8 / INT8 | ~43B parameters |
| 4-bit | ~87B parameters |
Rental cost
Indicative on-demand pricing runs roughly $0.70–$1.80 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.