RTX A6000
Ampere flagship professional GPU. 48GB GDDR6 with NVLink 3.0. Predecessor to RTX 6000 Ada. Widely available at lower cost.
RTX A6000 Overview
The RTX A6000 is an Ampere professional GPU with 48GB of ECC GDDR6, 77.4 TFLOPS of FP16/BF16 performance, and a 300W power envelope. As the predecessor to RTX 6000 Ada, it trades newer architecture performance for mature professional drivers, large memory, and broad availability.
Its 768 GB/s memory bandwidth is adequate for many workstation and inference workloads, while 48GB gives much more room than 24GB consumer cards for model weights, assets, and batch state. NVLink 3.0 allows paired professional cards to work together, but GDDR6 bandwidth remains a constraint compared with A100-class HBM platforms.
The RTX A6000 works well for professional rendering plus AI, larger-model inference, and cost-conscious multi-GPU workstation setups. It lacks FP8 support and does not match modern Ada or Hopper cards for pure inference throughput, so it is best chosen for capacity and professional features rather than speed alone.
Memory
Compute Performance
Hardware
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare RTX A6000 vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Ampere-class hardware.
Related GPUs
Frequently Asked Questions
How much VRAM does the RTX A6000 have?
The RTX A6000 has 48GB of GDDR6 memory with 768 GB/s bandwidth. This enables running models up to approximately 96B parameters at INT4 precision, 48B at INT8, or 24B at FP16.
What is the FP16 performance of the RTX A6000?
The RTX A6000 delivers 77.4 TFLOPS of FP16 performance and 77.4 TFLOPS BF16. INT8 throughput is 309 TOPS. For transformer inference, memory bandwidth (768 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the RTX A6000 best used for?
The RTX A6000 is best suited for: Large model inference, Multi-GPU NVLink setups, Professional rendering + AI. Ampere flagship professional GPU. 48GB GDDR6 with NVLink 3.0. Predecessor to RTX 6000 Ada. Widely available at lower cost.
What interconnect does the RTX A6000 use?
The RTX A6000 uses NVLink 3.0 / PCIe 4.0 with 112 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.
What LLM model sizes can the RTX A6000 run?
With 48GB of GDDR6, the RTX A6000 can run models up to approximately 24B parameters at FP16 (2 bytes/param), 48B at INT8 (1 byte/param), or 96B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the RTX A6000 compare to the A100 for LLM inference?
The RTX A6000 has 77.4 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 768 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the RTX A6000's lower cost.
What is the power consumption of the RTX A6000?
The RTX A6000 has a TDP (Thermal Design Power) of 300W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 300W, the RTX A6000 is in the mid-range tier — compatible with standard data center power infrastructure.
Ready to rent?
Compare RTX A6000 prices across 97+ providers
Live on-demand & spot rates · monthly cost estimates · availability status