Compute Comparison
NVIDIABlackwell2025

RTX 5080

Strong mid-range Blackwell. 16GB GDDR7 limits model size but excellent FP32 per dollar. Lower TDP than 5090.

VRAM
16GB
GDDR7
FP16
275.4
TFLOPS
Bandwidth
960.0
GB/s
TDP
360W
power
Best for:Budget consumer inferenceSmall model fine-tuningCost-efficient serving

RTX 5080 Overview

The RTX 5080 brings Blackwell architecture to a more attainable consumer tier, with 16GB of GDDR7, 275.4 TFLOPS of FP16/BF16 throughput, and a 360W board power target. It is a fast accelerator for compact models, but its memory capacity—not arithmetic capability—sets the practical boundary for many AI jobs.

GDDR7 delivers 960 GB/s of bandwidth, a substantial improvement for a 16GB-class card and useful for responsive inference on models that fit. Capacity remains the constraint: 16GB is comfortable for smaller models and quantized workloads, while larger FP16 models and large KV caches quickly require a different tier of hardware.

The RTX 5080 suits experimentation, small-model fine-tuning, image generation, and economical inference fleets where individual requests fit on one device. It lacks ECC and NVLink, so it is not a replacement for professional or HBM-equipped GPUs in reliability-critical or memory-scaled deployments.

Memory

VRAM16 GB
Memory TypeGDDR7
Bandwidth960 GB/s

Compute Performance

FP32137.7 TFLOPS
FP16275.4 TFLOPS
BF16275.4 TFLOPS
INT8551 TOPS

Hardware Specifications

Chip

ArchitectureGB203
GenerationBlackwell
Process NodeTSMC 4NP
Transistors45.6B
Die Size378 mm²
Release DateJanuary 30, 2025
Launch MSRP$999

Processors

CUDA / Shader Cores10,752
Tensor Cores336
RT Cores84

Clocks

Base Clock2,295 MHz
Boost Clock2,617 MHz

Memory

VRAM16 GB
Memory TypeGDDR7
Memory Bus256-bit
Bandwidth960 GB/s
L2 Cache64 MB

Power

TDP360 W
InterconnectPCIe 5.0

Relative Performance

FP16 Compute4%
VRAM Capacity6%
Mem Bandwidth6%

Relative to highest-spec GPU in database

Limitations

Only 16GB VRAM limits models to ~13B parameters at FP16
No NVLink — single-card VRAM ceiling
Consumer-grade reliability, no ECC memory

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX 5080 vs…

Use Case Guidance

Budget consumer inference
Small model fine-tuning
Cost-efficient serving

LLM Model Size Guidance

Max model (FP16)~8Bparameters at FP16 precision
Max model (INT8)~16Bparameters at INT8 precision
Max model (INT4)~32Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Blackwell-class hardware.

Browse all 42 LLM models

RTX 5080 vs Alternatives — Spec Comparison

SpecRTX 5080 thisRTX PRO 6000 BlackwellA100 80GBA100 40GB
VRAM16GB GDDR796GB GDDR780GB HBM2e40GB HBM2e
Memory Bandwidth960 GB/s1792 GB/s2000 GB/s1555 GB/s
FP16 TFLOPS275.4250312312
BF16 TFLOPS275.4250312312
FP8 TFLOPS500
INT8 TOPS551500624624
TDP360W300W400W300W
Process NodeTSMC 4NPTSMC 4NPTSMC 7nmTSMC 7nm
ArchitectureGB203GB202GA100GA100
Release Year2025202520202020
Max model (FP16)~8B params~48B params~40B params~20B params
Max model (INT4)~32B params~192B params~160B params~80B params
▲ indicates best value in row · FP16/BF16 TFLOPS at full precision · Max model estimates at 2 bytes/param (FP16) and 0.5 bytes/param (INT4)Full side-by-side comparison

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX 5080 have?

The RTX 5080 has 16GB of GDDR7 memory with 960 GB/s bandwidth. This enables running models up to approximately 32B parameters at INT4 precision, 16B at INT8, or 8B at FP16.

What is the FP16 performance of the RTX 5080?

The RTX 5080 delivers 275.4 TFLOPS of FP16 performance and 275.4 TFLOPS BF16. INT8 throughput is 551 TOPS. For transformer inference, memory bandwidth (960 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX 5080 best used for?

The RTX 5080 is best suited for: Budget consumer inference, Small model fine-tuning, Cost-efficient serving. Strong mid-range Blackwell. 16GB GDDR7 limits model size but excellent FP32 per dollar. Lower TDP than 5090.

What interconnect does the RTX 5080 use?

The RTX 5080 uses PCIe 5.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the RTX 5080 run?

With 16GB of GDDR7, the RTX 5080 can run models up to approximately 8B parameters at FP16 (2 bytes/param), 16B at INT8 (1 byte/param), or 32B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX 5080 compare to the A100 for LLM inference?

The RTX 5080 has 275.4 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 960 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX 5080's lower cost.

What is the power consumption of the RTX 5080?

The RTX 5080 has a TDP (Thermal Design Power) of 360W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 360W, the RTX 5080 is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare RTX 5080 prices across 102+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices