Compute Comparison
NVIDIAAmpere2021

A30

Slimmed-down A100 die with HBM2. Excellent memory bandwidth for its TDP class.

VRAM
24GB
HBM2
FP16
165.0
TFLOPS
Bandwidth
933.0
GB/s
TDP
165W
power
Best for:Inference at scaleLow-power HPCEdge data centers

A30 Overview

The NVIDIA A30 is a 165W Ampere data-center accelerator derived from the GA100 family, delivering 24GB of HBM2 and 165 TFLOPS of FP16/BF16 throughput. It targets deployments that need data-center features and HBM bandwidth without the power draw or cost of a full A100.

Its 933 GB/s of HBM2 bandwidth is unusually strong for a 24GB, low-power card and is valuable for serving models that fit in memory. The 24GB capacity still limits full-precision model placement and context cache, but NVLink 3.0 and MIG support give operators options for multi-GPU layouts and partitioned, multi-tenant services.

The A30 is a sensible fit for efficient inference at scale, lower-power HPC, and organizations that need predictable Ampere behavior. It is not ideal for very large models, FP8 pipelines, or tasks that require high FP32 performance for graphics-heavy or simulation-oriented work.

Memory

VRAM24 GB
Memory TypeHBM2
Bandwidth933 GB/s
NVLink BW200 GB/s

Compute Performance

FP3210.3 TFLOPS
FP16165 TFLOPS
BF16165 TFLOPS
INT8330 TOPS

Hardware

ArchitectureGA100
GenerationAmpere
Process NodeTSMC 7nm
Transistors54.2B
TDP165 W
InterconnectNVLink 3.0 / PCIe 4.0
Release Year2021

Relative Performance

FP16 Compute2%
VRAM Capacity8%
Mem Bandwidth6%

Relative to highest-spec GPU in database

Limitations

Only 24GB HBM2 — limits to ~13B models at FP16
Low FP32 throughput (10.3 TFLOPS) vs newer GPUs
Limited cloud availability compared to A100/H100

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare A30 vs…

Use Case Guidance

Inference at scale
Low-power HPC
Edge data centers

LLM Model Size Guidance

Max model (FP16)~12Bparameters at FP16 precision
Max model (INT8)~24Bparameters at INT8 precision
Max model (INT4)~48Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ampere-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the A30 have?

The A30 has 24GB of HBM2 memory with 933 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.

What is the FP16 performance of the A30?

The A30 delivers 165 TFLOPS of FP16 performance and 165 TFLOPS BF16. INT8 throughput is 330 TOPS. For transformer inference, memory bandwidth (933 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the A30 best used for?

The A30 is best suited for: Inference at scale, Low-power HPC, Edge data centers. Slimmed-down A100 die with HBM2. Excellent memory bandwidth for its TDP class.

What interconnect does the A30 use?

The A30 uses NVLink 3.0 / PCIe 4.0 with 200 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.

What LLM model sizes can the A30 run?

With 24GB of HBM2, the A30 can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the A30 compare to the A100 for LLM inference?

The A30 has 165 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 933 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the A30's lower cost.

What is the power consumption of the A30?

The A30 has a TDP (Thermal Design Power) of 165W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 165W, the A30 is in the low-power tier — enables high-density deployments with standard rack power.

Ready to rent?

Compare A30 prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices