Compute Comparison
NVIDIAAmpere2020

RTX A6000

Ampere flagship professional GPU. 48GB GDDR6 with NVLink 3.0. Predecessor to RTX 6000 Ada. Widely available at lower cost.

VRAM
48GB
GDDR6
FP16
77.4
TFLOPS
Bandwidth
768.0
GB/s
TDP
300W
power
Best for:Large model inferenceMulti-GPU NVLink setupsProfessional rendering + AI

RTX A6000 Overview

The RTX A6000 is an Ampere professional GPU with 48GB of ECC GDDR6, 77.4 TFLOPS of FP16/BF16 performance, and a 300W power envelope. As the predecessor to RTX 6000 Ada, it trades newer architecture performance for mature professional drivers, large memory, and broad availability.

Its 768 GB/s memory bandwidth is adequate for many workstation and inference workloads, while 48GB gives much more room than 24GB consumer cards for model weights, assets, and batch state. NVLink 3.0 allows paired professional cards to work together, but GDDR6 bandwidth remains a constraint compared with A100-class HBM platforms.

The RTX A6000 works well for professional rendering plus AI, larger-model inference, and cost-conscious multi-GPU workstation setups. It lacks FP8 support and does not match modern Ada or Hopper cards for pure inference throughput, so it is best chosen for capacity and professional features rather than speed alone.

Memory

VRAM48 GB
Memory TypeGDDR6
Bandwidth768 GB/s
NVLink BW112 GB/s

Compute Performance

FP3238.7 TFLOPS
FP1677.4 TFLOPS
BF1677.4 TFLOPS
INT8309 TOPS

Hardware

ArchitectureGA102
GenerationAmpere
Process NodeSamsung 8nm
Transistors28.3B
TDP300 W
InterconnectNVLink 3.0 / PCIe 4.0
Release Year2020

Relative Performance

FP16 Compute1%
VRAM Capacity17%
Mem Bandwidth5%

Relative to highest-spec GPU in database

Limitations

GDDR6 memory bandwidth far below HBM alternatives
No NVLink on A4000/A2000 — single-card VRAM ceiling
Professional but older Ampere architecture — no FP8 support

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare RTX A6000 vs…

Use Case Guidance

Large model inference
Multi-GPU NVLink setups
Professional rendering + AI

LLM Model Size Guidance

Max model (FP16)~24Bparameters at FP16 precision
Max model (INT8)~48Bparameters at INT8 precision
Max model (INT4)~96Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ampere-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the RTX A6000 have?

The RTX A6000 has 48GB of GDDR6 memory with 768 GB/s bandwidth. This enables running models up to approximately 96B parameters at INT4 precision, 48B at INT8, or 24B at FP16.

What is the FP16 performance of the RTX A6000?

The RTX A6000 delivers 77.4 TFLOPS of FP16 performance and 77.4 TFLOPS BF16. INT8 throughput is 309 TOPS. For transformer inference, memory bandwidth (768 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the RTX A6000 best used for?

The RTX A6000 is best suited for: Large model inference, Multi-GPU NVLink setups, Professional rendering + AI. Ampere flagship professional GPU. 48GB GDDR6 with NVLink 3.0. Predecessor to RTX 6000 Ada. Widely available at lower cost.

What interconnect does the RTX A6000 use?

The RTX A6000 uses NVLink 3.0 / PCIe 4.0 with 112 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.

What LLM model sizes can the RTX A6000 run?

With 48GB of GDDR6, the RTX A6000 can run models up to approximately 24B parameters at FP16 (2 bytes/param), 48B at INT8 (1 byte/param), or 96B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the RTX A6000 compare to the A100 for LLM inference?

The RTX A6000 has 77.4 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 768 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX A6000's lower cost.

What is the power consumption of the RTX A6000?

The RTX A6000 has a TDP (Thermal Design Power) of 300W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 300W, the RTX A6000 is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare RTX A6000 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices