RTX 5080
Strong mid-range Blackwell. 16GB GDDR7 limits model size but excellent FP32 per dollar. Lower TDP than 5090.
RTX 5080 Overview
The RTX 5080 brings Blackwell architecture to a more attainable consumer tier, with 16GB of GDDR7, 275.4 TFLOPS of FP16/BF16 throughput, and a 360W board power target. It is a fast accelerator for compact models, but its memory capacity—not arithmetic capability—sets the practical boundary for many AI jobs.
GDDR7 delivers 960 GB/s of bandwidth, a substantial improvement for a 16GB-class card and useful for responsive inference on models that fit. Capacity remains the constraint: 16GB is comfortable for smaller models and quantized workloads, while larger FP16 models and large KV caches quickly require a different tier of hardware.
The RTX 5080 suits experimentation, small-model fine-tuning, image generation, and economical inference fleets where individual requests fit on one device. It lacks ECC and NVLink, so it is not a replacement for professional or HBM-equipped GPUs in reliability-critical or memory-scaled deployments.
Memory
Compute Performance
Hardware
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare RTX 5080 vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Blackwell-class hardware.
Related GPUs
Frequently Asked Questions
How much VRAM does the RTX 5080 have?
The RTX 5080 has 16GB of GDDR7 memory with 960 GB/s bandwidth. This enables running models up to approximately 32B parameters at INT4 precision, 16B at INT8, or 8B at FP16.
What is the FP16 performance of the RTX 5080?
The RTX 5080 delivers 275.4 TFLOPS of FP16 performance and 275.4 TFLOPS BF16. INT8 throughput is 551 TOPS. For transformer inference, memory bandwidth (960 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the RTX 5080 best used for?
The RTX 5080 is best suited for: Budget consumer inference, Small model fine-tuning, Cost-efficient serving. Strong mid-range Blackwell. 16GB GDDR7 limits model size but excellent FP32 per dollar. Lower TDP than 5090.
What interconnect does the RTX 5080 use?
The RTX 5080 uses PCIe 5.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.
What LLM model sizes can the RTX 5080 run?
With 16GB of GDDR7, the RTX 5080 can run models up to approximately 8B parameters at FP16 (2 bytes/param), 16B at INT8 (1 byte/param), or 32B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the RTX 5080 compare to the A100 for LLM inference?
The RTX 5080 has 275.4 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 960 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 5080's lower cost.
What is the power consumption of the RTX 5080?
The RTX 5080 has a TDP (Thermal Design Power) of 360W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 360W, the RTX 5080 is in the mid-range tier — compatible with standard data center power infrastructure.
Ready to rent?
Compare RTX 5080 prices across 97+ providers
Live on-demand & spot rates · monthly cost estimates · availability status