A30
Slimmed-down A100 die with HBM2. Excellent memory bandwidth for its TDP class.
A30 Overview
The NVIDIA A30 is a 165W Ampere data-center accelerator derived from the GA100 family, delivering 24GB of HBM2 and 165 TFLOPS of FP16/BF16 throughput. It targets deployments that need data-center features and HBM bandwidth without the power draw or cost of a full A100.
Its 933 GB/s of HBM2 bandwidth is unusually strong for a 24GB, low-power card and is valuable for serving models that fit in memory. The 24GB capacity still limits full-precision model placement and context cache, but NVLink 3.0 and MIG support give operators options for multi-GPU layouts and partitioned, multi-tenant services.
The A30 is a sensible fit for efficient inference at scale, lower-power HPC, and organizations that need predictable Ampere behavior. It is not ideal for very large models, FP8 pipelines, or tasks that require high FP32 performance for graphics-heavy or simulation-oriented work.
Memory
Compute Performance
Hardware
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare A30 vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Ampere-class hardware.
Related GPUs
Frequently Asked Questions
How much VRAM does the A30 have?
The A30 has 24GB of HBM2 memory with 933 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.
What is the FP16 performance of the A30?
The A30 delivers 165 TFLOPS of FP16 performance and 165 TFLOPS BF16. INT8 throughput is 330 TOPS. For transformer inference, memory bandwidth (933 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the A30 best used for?
The A30 is best suited for: Inference at scale, Low-power HPC, Edge data centers. Slimmed-down A100 die with HBM2. Excellent memory bandwidth for its TDP class.
What interconnect does the A30 use?
The A30 uses NVLink 3.0 / PCIe 4.0 with 200 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.
What LLM model sizes can the A30 run?
With 24GB of HBM2, the A30 can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the A30 compare to the A100 for LLM inference?
The A30 has 165 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 933 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the A30's lower cost.
What is the power consumption of the A30?
The A30 has a TDP (Thermal Design Power) of 165W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 165W, the A30 is in the low-power tier — enables high-density deployments with standard rack power.
Ready to rent?
Compare A30 prices across 97+ providers
Live on-demand & spot rates · monthly cost estimates · availability status