NVIDIA GeForce RTX 4090: specs, pricing and what it can run

Still the price/performance reference for local LLM work, though 24GB caps it at ~30B parameters in 4-bit.

Specifications

SpecNVIDIA GeForce RTX 4090
ArchitectureAda Lovelace
Released2022
Memory24 GB GDDR6X
Memory bandwidth1,008 GB/s
Dense FP16330 TFLOPS
FP8 tensor coresYes
Board power450 W
InterconnectPCIe Gen4 (no NVLink)
Segmentconsumer

What fits in 24 GB

Weights only, at 85% of nominal capacity to leave room for the KV cache and activations. Long-context serving needs materially more headroom than this table implies.

QuantisationLargest model (weights only)
FP16~10B parameters
FP8 / INT8~21B parameters
4-bit~43B parameters

Rental cost

Indicative on-demand pricing runs roughly $0.30–$0.90 per GPU-hour depending on provider, region, and commitment. Spot and reserved capacity sit well below that band. These are ranges rather than quotes: street prices move week to week, so check retail + live cloud price feed before budgeting.

Related

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.