RTX 5090
Flagship consumer Blackwell GPU. Massive FP32 uplift over 4090. 32GB GDDR7 enables larger models. Best consumer GPU for AI workloads.
RTX 5090 Overview
The RTX 5090 is NVIDIA’s flagship consumer Blackwell GPU, pairing 32GB of GDDR7 with 419.6 TFLOPS of FP16/BF16 tensor throughput. Its PCIe 5.0 design and 575W power envelope put it in a very different class from prior consumer cards, but it remains a consumer product rather than a data-center accelerator.
Its 1,792 GB/s memory bandwidth is exceptionally high for a non-HBM card and helps feed inference workloads that would otherwise be bandwidth-bound. The 32GB capacity is enough for many 13B to 30B-class workflows with careful precision and batch-size choices, but it is still a fixed single-card memory ceiling because there is no NVLink or ECC protection.
This is a strong fit for high-throughput local inference, fine-tuning that fits in 32GB, and development environments where consumer economics are attractive. It is less appropriate for unattended, compliance-sensitive production infrastructure, very large FP16 models, or workloads that need pooled multi-GPU memory.
Memory
Compute Performance
Hardware
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare RTX 5090 vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Blackwell-class hardware.
Related GPUs
Frequently Asked Questions
How much VRAM does the RTX 5090 have?
The RTX 5090 has 32GB of GDDR7 memory with 1792 GB/s bandwidth. This enables running models up to approximately 64B parameters at INT4 precision, 32B at INT8, or 16B at FP16.
What is the FP16 performance of the RTX 5090?
The RTX 5090 delivers 419.6 TFLOPS of FP16 performance and 419.6 TFLOPS BF16. INT8 throughput is 839 TOPS. For transformer inference, memory bandwidth (1792 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the RTX 5090 best used for?
The RTX 5090 is best suited for: Consumer AI inference, Fine-tuning up to 30B, High-throughput serving. Flagship consumer Blackwell GPU. Massive FP32 uplift over 4090. 32GB GDDR7 enables larger models. Best consumer GPU for AI workloads.
What interconnect does the RTX 5090 use?
The RTX 5090 uses PCIe 5.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.
What LLM model sizes can the RTX 5090 run?
With 32GB of GDDR7, the RTX 5090 can run models up to approximately 16B parameters at FP16 (2 bytes/param), 32B at INT8 (1 byte/param), or 64B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the RTX 5090 compare to the A100 for LLM inference?
The RTX 5090 has 419.6 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 1792 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 5090's lower cost.
What is the power consumption of the RTX 5090?
The RTX 5090 has a TDP (Thermal Design Power) of 575W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 575W, the RTX 5090 is in the high-power tier — requires specialized data center infrastructure with high-density power delivery.
Ready to rent?
Compare RTX 5090 prices across 97+ providers
Live on-demand & spot rates · monthly cost estimates · availability status