Compute Comparison
NVIDIABlackwell2025

RTX 5090

Flagship consumer Blackwell GPU. Massive FP32 uplift over 4090. 32GB GDDR7 enables larger models. Best consumer GPU for AI workloads.

VRAM
32GB
GDDR7
FP16
419.6
TFLOPS
Bandwidth
1.8k
GB/s
TDP
575W
power
Best for:Consumer AI inferenceFine-tuning up to 30BHigh-throughput serving

RTX 5090 Overview

The RTX 5090 is NVIDIA’s flagship consumer Blackwell GPU, pairing 32GB of GDDR7 with 419.6 TFLOPS of FP16/BF16 tensor throughput. Its PCIe 5.0 design and 575W power envelope put it in a very different class from prior consumer cards, but it remains a consumer product rather than a data-center accelerator.

Its 1,792 GB/s memory bandwidth is exceptionally high for a non-HBM card and helps feed inference workloads that would otherwise be bandwidth-bound. The 32GB capacity is enough for many 13B to 30B-class workflows with careful precision and batch-size choices, but it is still a fixed single-card memory ceiling because there is no NVLink or ECC protection.

This is a strong fit for high-throughput local inference, fine-tuning that fits in 32GB, and development environments where consumer economics are attractive. It is less appropriate for unattended, compliance-sensitive production infrastructure, very large FP16 models, or workloads that need pooled multi-GPU memory.

Memory

VRAM32 GB
Memory TypeGDDR7
Bandwidth1792 GB/s

Compute Performance

FP32209.8 TFLOPS
FP16419.6 TFLOPS
BF16419.6 TFLOPS
INT8839 TOPS

Hardware

ArchitectureGB202
GenerationBlackwell
Process NodeTSMC 4NP
Transistors92.2B
TDP575 W
InterconnectPCIe 5.0
Release Year2025

Relative Performance

FP16 Compute6%
VRAM Capacity11%
Mem Bandwidth11%

Relative to highest-spec GPU in database

Limitations

Consumer GPU — no ECC memory, not suited for production data centers
No NVLink — cannot pool VRAM across cards
575W TDP requires high-end cooling infrastructure

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX 5090 vs…

Use Case Guidance

Consumer AI inference
Fine-tuning up to 30B
High-throughput serving

LLM Model Size Guidance

Max model (FP16)~16Bparameters at FP16 precision
Max model (INT8)~32Bparameters at INT8 precision
Max model (INT4)~64Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Blackwell-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX 5090 have?

The RTX 5090 has 32GB of GDDR7 memory with 1792 GB/s bandwidth. This enables running models up to approximately 64B parameters at INT4 precision, 32B at INT8, or 16B at FP16.

What is the FP16 performance of the RTX 5090?

The RTX 5090 delivers 419.6 TFLOPS of FP16 performance and 419.6 TFLOPS BF16. INT8 throughput is 839 TOPS. For transformer inference, memory bandwidth (1792 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX 5090 best used for?

The RTX 5090 is best suited for: Consumer AI inference, Fine-tuning up to 30B, High-throughput serving. Flagship consumer Blackwell GPU. Massive FP32 uplift over 4090. 32GB GDDR7 enables larger models. Best consumer GPU for AI workloads.

What interconnect does the RTX 5090 use?

The RTX 5090 uses PCIe 5.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the RTX 5090 run?

With 32GB of GDDR7, the RTX 5090 can run models up to approximately 16B parameters at FP16 (2 bytes/param), 32B at INT8 (1 byte/param), or 64B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX 5090 compare to the A100 for LLM inference?

The RTX 5090 has 419.6 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 1792 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 5090's lower cost.

What is the power consumption of the RTX 5090?

The RTX 5090 has a TDP (Thermal Design Power) of 575W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 575W, the RTX 5090 is in the high-power tier — requires specialized data center infrastructure with high-density power delivery.

Ready to rent?

Compare RTX 5090 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices