NVIDIA GeForce RTX 5090: specs, pricing and what it can run

The strongest consumer option for local inference: 32GB of GDDR7 at 1.79 TB/s beats every previous consumer part. Licence terms restrict datacenter deployment.

Specifications

SpecNVIDIA GeForce RTX 5090
ArchitectureBlackwell
Released2025
Memory32 GB GDDR7
Memory bandwidth1,792 GB/s
Dense FP16419 TFLOPS
FP8 tensor coresYes
Board power575 W
InterconnectPCIe Gen5 (no NVLink)
Segmentconsumer

What fits in 32 GB

Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.

QuantisationLargest model (weights only)
FP16~14B parameters
FP8 / INT8~29B parameters
4-bit~58B parameters

Rental cost

Indicative on-demand pricing runs roughly $0.40–$1.20 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check retail + live cloud price feed before budgeting.

Related

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.