NVIDIA L40S: specs, pricing and what it can run

Cost-efficient inference and fine-tuning part. GDDR6 bandwidth is a quarter of H100, so it suits smaller models and batch work rather than latency-critical serving.

Specifications

SpecNVIDIA L40S
ArchitectureAda Lovelace
Released2023
Memory48 GB GDDR6
Memory bandwidth864 GB/s
Dense FP16362 TFLOPS
FP8 tensor coresYes
Board power350 W
InterconnectPCIe Gen4 (no NVLink)
Segmentdatacenter

What fits in 48 GB

Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.

QuantisationLargest model (weights only)
FP16~21B parameters
FP8 / INT8~43B parameters
4-bit~87B parameters

Rental cost

Indicative on-demand pricing runs roughly $0.70–$1.80 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check live cloud price feed before budgeting.

Related

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.