A10G
Data center variant of RTX 3090 die. Very low TDP (150W). Ideal for inference at scale.
A10G Overview
The NVIDIA A10G is an Ampere-generation data center GPU designed for inference serving, built on the GA102 die (Samsung 8nm). At just 150W TDP, it is one of the most power-efficient data center GPUs available, enabling dense rack deployments without specialized cooling. It delivers 125 TFLOPS of FP16 throughput and 250 TOPS of INT8, making it well-suited for serving quantized models at scale. It is the primary GPU in AWS g5 instances.
The 24GB GDDR6 configuration provides 600 GB/s of memory bandwidth — lower than HBM-based alternatives but higher than the T4's 320 GB/s. This fits models up to approximately 12B parameters at FP16, or 24B at INT8. The A10G uses PCIe 4.0 with no NVLink, so VRAM cannot be pooled across cards. Its 31.2 TFLOPS FP32 is notably high for a data center GPU, making it well-suited for mixed AI inference and graphics rendering workloads.
The A10G is the right choice for inference serving of models up to 13B parameters, particularly on AWS where it is the primary GPU in g5 instances. It is well-suited for Stable Diffusion and image generation workloads, mixed compute and graphics environments, and high-density inference deployments where power efficiency matters. It is not suitable for training large models (limited VRAM, no NVLink) or for memory-bound inference of 70B+ models where HBM bandwidth matters.
Memory
Compute Performance
Hardware
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare A10G vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Ampere-class hardware.
Related GPUs
Frequently Asked Questions
How much VRAM does the A10G have?
The A10G has 24GB of GDDR6 memory with 600 GB/s bandwidth. This enables running models up to approximately 48B parameters at INT4 precision, 24B at INT8, or 12B at FP16.
What is the FP16 performance of the A10G?
The A10G delivers 125 TFLOPS of FP16 performance and 125 TFLOPS BF16. INT8 throughput is 250 TOPS. For transformer inference, memory bandwidth (600 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the A10G best used for?
The A10G is best suited for: Inference serving, Graphics workloads, Low-power deployments. Data center variant of RTX 3090 die. Very low TDP (150W). Ideal for inference at scale.
What interconnect does the A10G use?
The A10G uses PCIe 4.0. Without NVLink, VRAM cannot be pooled across multiple cards — the single-card capacity is the hard ceiling for model size.
What LLM model sizes can the A10G run?
With 24GB of GDDR6, the A10G can run models up to approximately 12B parameters at FP16 (2 bytes/param), 24B at INT8 (1 byte/param), or 48B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the A10G compare to the A100 for LLM inference?
The A10G has 125 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 600 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the A10G's lower cost.
What is the power consumption of the A10G?
The A10G has a TDP (Thermal Design Power) of 150W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 150W, the A10G is in the low-power tier — enables high-density deployments with standard rack power.
Ready to rent?
Compare A10G prices across 97+ providers
Live on-demand & spot rates · monthly cost estimates · availability status