Compute Comparison
NVIDIABlackwell2025

RTX 5080

Strong mid-range Blackwell. 16GB GDDR7 limits model size but excellent FP32 per dollar. Lower TDP than 5090.

VRAM
16GB
GDDR7
FP16
275.4
TFLOPS
Bandwidth
960.0
GB/s
TDP
360W
power
Best for:Budget consumer inferenceSmall model fine-tuningCost-efficient serving

RTX 5080 Overview

The RTX 5080 brings Blackwell architecture to a more attainable consumer tier, with 16GB of GDDR7, 275.4 TFLOPS of FP16/BF16 throughput, and a 360W board power target. It is a fast accelerator for compact models, but its memory capacity—not arithmetic capability—sets the practical boundary for many AI jobs.

GDDR7 delivers 960 GB/s of bandwidth, a substantial improvement for a 16GB-class card and useful for responsive inference on models that fit. Capacity remains the constraint: 16GB is comfortable for smaller models and quantized workloads, while larger FP16 models and large KV caches quickly require a different tier of hardware.

The RTX 5080 suits experimentation, small-model fine-tuning, image generation, and economical inference fleets where individual requests fit on one device. It lacks ECC and NVLink, so it is not a replacement for professional or HBM-equipped GPUs in reliability-critical or memory-scaled deployments.

Memory

VRAM16 GB
Memory TypeGDDR7
Bandwidth960 GB/s

Compute Performance

FP32137.7 TFLOPS
FP16275.4 TFLOPS
BF16275.4 TFLOPS
INT8551 TOPS

Hardware

ArchitectureGB203
GenerationBlackwell
Process NodeTSMC 4NP
Transistors45.6B
TDP360 W
InterconnectPCIe 5.0
Release Year2025

Relative Performance

FP16 Compute4%
VRAM Capacity6%
Mem Bandwidth6%

Relative to highest-spec GPU in database

Limitations

Only 16GB VRAM limits models to ~13B parameters at FP16
No NVLink — single-card VRAM ceiling
Consumer-grade reliability, no ECC memory

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX 5080 vs…

Use Case Guidance

Budget consumer inference
Small model fine-tuning
Cost-efficient serving

LLM Model Size Guidance

Max model (FP16)~8Bparameters at FP16 precision
Max model (INT8)~16Bparameters at INT8 precision
Max model (INT4)~32Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Blackwell-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX 5080 have?

The RTX 5080 has 16GB of GDDR7 memory with 960 GB/s bandwidth. This enables running models up to approximately 32B parameters at INT4 precision, 16B at INT8, or 8B at FP16.

What is the FP16 performance of the RTX 5080?

The RTX 5080 delivers 275.4 TFLOPS of FP16 performance and 275.4 TFLOPS BF16. INT8 throughput is 551 TOPS. For transformer inference, memory bandwidth (960 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX 5080 best used for?

The RTX 5080 is best suited for: Budget consumer inference, Small model fine-tuning, Cost-efficient serving. Strong mid-range Blackwell. 16GB GDDR7 limits model size but excellent FP32 per dollar. Lower TDP than 5090.

What interconnect does the RTX 5080 use?

The RTX 5080 uses PCIe 5.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the RTX 5080 run?

With 16GB of GDDR7, the RTX 5080 can run models up to approximately 8B parameters at FP16 (2 bytes/param), 16B at INT8 (1 byte/param), or 32B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX 5080 compare to the A100 for LLM inference?

The RTX 5080 has 275.4 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 960 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 5080's lower cost.

What is the power consumption of the RTX 5080?

The RTX 5080 has a TDP (Thermal Design Power) of 360W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 360W, the RTX 5080 is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare RTX 5080 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices