Compute Comparison
NVIDIAAda Lovelace2022

RTX 4090

Consumer GPU with surprisingly strong FP32. No NVLink. Best $/TFLOP for budget workloads.

VRAM
24GB
GDDR6X
FP16
165.2
TFLOPS
Bandwidth
1.0k
GB/s
TDP
450W
power
Best for:Fine-tuning small modelsInference up to 13BCost-sensitive workloads

RTX 4090 Overview

The NVIDIA RTX 4090 is the most powerful consumer GPU available and the top choice for budget-conscious AI workloads that fit within 24GB of VRAM. Built on the Ada Lovelace architecture (AD102 die, TSMC 4N process), it delivers 165 TFLOPS of FP16 throughput — comparable to some data center GPUs at a fraction of the cost. Its 82.6 TFLOPS FP32 figure is particularly strong, making it well-suited for Stable Diffusion and other FP32-heavy generative AI workloads.

The 24GB GDDR6X configuration provides 1,008 GB/s of memory bandwidth — higher than the A100 40GB's 1,555 GB/s but lower than HBM-based data center GPUs. This fits models up to approximately 12B parameters at FP16, or up to 48B parameters at INT4/GGUF quantization. The RTX 4090 lacks NVLink, so VRAM cannot be pooled across cards. It also lacks ECC memory, making it unsuitable for production data center deployments where memory error correction is required.

The RTX 4090 is the right choice for fine-tuning models up to 13B parameters, running inference on quantized models up to 70B (INT4), and for Stable Diffusion and image generation workloads. It is not appropriate for production data center use (no ECC), training models above 13B parameters, or workloads requiring multi-GPU VRAM pooling. At $0.35–$0.75/hr on cloud providers, it offers the best cost-per-TFLOP of any widely available GPU — making it the default choice for experimentation and development.

Memory

VRAM24 GB
Memory TypeGDDR6X
Bandwidth1008 GB/s

Compute Performance

FP3282.6 TFLOPS
FP16165.2 TFLOPS
BF16165.2 TFLOPS
INT8330.3 TOPS

Hardware

ArchitectureAD102
GenerationAda Lovelace
Process NodeTSMC 4N
Transistors76.3B
TDP450 W
InterconnectPCIe 4.0
Release Year2022

Relative Performance

FP16 Compute2%
VRAM Capacity8%
Mem Bandwidth6%

Relative to highest-spec GPU in database

Limitations

Consumer GPU — no ECC memory, not suited for production data centers
No NVLink — cannot pool VRAM across cards
Only 24GB VRAM — limits to ~13B models at FP16

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX 4090 vs…

Use Case Guidance

Fine-tuning small models
Inference up to 13B
Cost-sensitive workloads

LLM Model Size Guidance

Max model (FP16)~12Bparameters at FP16 precision
Max model (INT8)~24Bparameters at INT8 precision
Max model (INT4)~48Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ada Lovelace-class hardware.

Browse all 42 LLM models Self-hosting guide: best LLM models for this GPU

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX 4090 have?

The RTX 4090 has 24GB of GDDR6X memory with 1008 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.

What is the FP16 performance of the RTX 4090?

The RTX 4090 delivers 165.2 TFLOPS of FP16 performance and 165.2 TFLOPS BF16. INT8 throughput is 330.3 TOPS. For transformer inference, memory bandwidth (1008 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX 4090 best used for?

The RTX 4090 is best suited for: Fine-tuning small models, Inference up to 13B, Cost-sensitive workloads. Consumer GPU with surprisingly strong FP32. No NVLink. Best $/TFLOP for budget workloads.

What interconnect does the RTX 4090 use?

The RTX 4090 uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the RTX 4090 run?

With 24GB of GDDR6X, the RTX 4090 can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX 4090 compare to the A100 for LLM inference?

The RTX 4090 has 165.2 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 1008 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 4090's lower cost.

What is the power consumption of the RTX 4090?

The RTX 4090 has a TDP (Thermal Design Power) of 450W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 450W, the RTX 4090 is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare RTX 4090 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices