Compute Comparison
NVIDIAAmpere2021

A10G

Data center variant of RTX 3090 die. Very low TDP (150W). Ideal for inference at scale.

VRAM
24GB
GDDR6
FP16
125.0
TFLOPS
Bandwidth
600.0
GB/s
TDP
150W
power
Best for:Inference servingGraphics workloadsLow-power deployments

A10G Overview

The NVIDIA A10G is an Ampere-generation data center GPU designed for inference serving, built on the GA102 die (Samsung 8nm). At just 150W TDP, it is one of the most power-efficient data center GPUs available, enabling dense rack deployments without specialized cooling. It delivers 125 TFLOPS of FP16 throughput and 250 TOPS of INT8, making it well-suited for serving quantized models at scale. It is the primary GPU in AWS g5 instances.

The 24GB GDDR6 configuration provides 600 GB/s of memory bandwidth — lower than HBM-based alternatives but higher than the T4's 320 GB/s. This fits models up to approximately 12B parameters at FP16, or 24B at INT8. The A10G uses PCIe 4.0 with no NVLink, so VRAM cannot be pooled across cards. Its 31.2 TFLOPS FP32 is notably high for a data center GPU, making it well-suited for mixed AI inference and graphics rendering workloads.

The A10G is the right choice for inference serving of models up to 13B parameters, particularly on AWS where it is the primary GPU in g5 instances. It is well-suited for Stable Diffusion and image generation workloads, mixed compute and graphics environments, and high-density inference deployments where power efficiency matters. It is not suitable for training large models (limited VRAM, no NVLink) or for memory-bound inference of 70B+ models where HBM bandwidth matters.

Memory

VRAM24 GB
Memory TypeGDDR6
Bandwidth600 GB/s

Compute Performance

FP3231.2 TFLOPS
FP16125 TFLOPS
BF16125 TFLOPS
INT8250 TOPS

Hardware

ArchitectureGA102
GenerationAmpere
Process NodeSamsung 8nm
Transistors28.3B
TDP150 W
InterconnectPCIe 4.0
Release Year2021

Relative Performance

FP16 Compute2%
VRAM Capacity8%
Mem Bandwidth4%

Relative to highest-spec GPU in database

Limitations

Only 24GB GDDR6 — limits to ~13B models at FP16
No NVLink — single-card VRAM ceiling
Lower memory bandwidth than HBM-based alternatives

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare A10G vs…

Use Case Guidance

Inference serving
Graphics workloads
Low-power deployments

LLM Model Size Guidance

Max model (FP16)~12Bparameters at FP16 precision
Max model (INT8)~24Bparameters at INT8 precision
Max model (INT4)~48Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Ampere-class hardware.

Browse all 42 LLM models Self-hosting guide: best LLM models for this GPU

Related GPUs

Frequently Asked Questions

How much VRAM does the A10G have?

The A10G has 24GB of GDDR6 memory with 600 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.

What is the FP16 performance of the A10G?

The A10G delivers 125 TFLOPS of FP16 performance and 125 TFLOPS BF16. INT8 throughput is 250 TOPS. For transformer inference, memory bandwidth (600 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the A10G best used for?

The A10G is best suited for: Inference serving, Graphics workloads, Low-power deployments. Data center variant of RTX 3090 die. Very low TDP (150W). Ideal for inference at scale.

What interconnect does the A10G use?

The A10G uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.

What LLM model sizes can the A10G run?

With 24GB of GDDR6, the A10G can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the A10G compare to the A100 for LLM inference?

The A10G has 125 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 600 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the A10G's lower cost.

What is the power consumption of the A10G?

The A10G has a TDP (Thermal Design Power) of 150W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 150W, the A10G is in the low-power tier — enables high-density deployments with standard rack power.

Ready to rent?

Compare A10G prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices