GPU Specs Database — H100, A100, RTX 4090 & All AI Cloud GPU Specs
Complete technical specifications for every major AI and cloud GPU — NVIDIA Hopper, Blackwell, Ampere, Ada Lovelace, AMD CDNA 3, and more. Click any GPU to see live cloud pricing across 94+ providers.
GB300 NVL72
GB200 NVL72
B200 192GB
B200 SXM 192GB
B200 NVL 192GB
B300 SXM6
GB300 SXM6
B300 SXM 262GB
GB300 SXM 288GB
H200 141GB
H100 80GB
H100 SXM5 80GB
H100 NVL 94GB
H200 SXM 141GB
H200 NVL 141GB
GH200 Grace Hopper 96GB
H100 40GB
H100 PCIe 80GB
MI325X 256GB
MI325X 256GB
MI300X 192GB
MI300X 192GB
MI355X
MI355X 288GB
Gaudi 2 96GB
Gaudi 2
RTX 5090
MI250X 128GB
A100 80GB
A100 40GB
A100 SXM4 80GB
A100 PCIe 80GB
A100 PCIe 40GB
A100 SXM4 40GB
RTX 5080
RTX PRO 6000 Blackwell
RTX 5070 Ti
L40S
RTX 6000 Ada
L40
RTX 5880 Ada
RTX 4090
A30
RTX 5070
RTX 5000 Ada
Quadro RTX 6000 24GB
Quadro RTX 8000 48GB
A10G
A10 24GB
V100 32GB
V100 16GB
TITAN V 12GB
RTX 4080
RTX 4500 Ada
Quadro RTX 5000 16GB
RTX 4070 Ti
RTX A6000
A40
RTX 3090
RTX 5060 Ti
RTX 3080 Ti
RTX A5500 24GB
T4
A16
L4
RTX 3080
RTX 4070
Quadro RTX 4000 8GB
RTX A5000
RTX 4000 Ada
RTX 5060
RTX A4500 20GB
RTX 4060 Ti 16GB
RTX 3070 8GB
RTX 4000 SFF Ada
RTX A4000 16GB
RTX 3060 Ti 8GB
RTX 4060 8GB
RTX 3060 12GB
RTX 2000 Ada
Tesla P100 16GB
RTX A2000 12GB
TITAN Xp 12GB
Quadro P6000 24GB
Quadro P5000 16GB
Quadro P4000 8GB
Quadro M4000 8GB
GPU specs guide for AI and cloud compute
NVIDIA Hopper: H100 and H200
NVIDIA's Hopper architecture (2022–2023) remains the dominant GPU for large-scale AI training and inference in 2026. The H100 SXM delivers 989 TFLOPS FP16 with 80GB HBM3 at 3.35 TB/s bandwidth and 900 GB/s NVLink 4.0 interconnect — the standard for 70B+ LLM training. The H100 PCIe variant offers the same compute at lower bandwidth (2 TB/s) and is 10–20% cheaper on cloud. The H200 upgrades memory to 141GB HBM3e at 4.8 TB/s bandwidth — a significant advantage for inference serving of large models that previously required multi-GPU configurations. H100 on-demand pricing ranges from $3.50–$5.50/hr depending on provider.
NVIDIA Blackwell: B200 and GB200
NVIDIA's Blackwell architecture (2024–2025) delivers a 2.5–4× performance improvement over Hopper for AI inference. The B200 provides 180GB HBM3e memory — enough to serve a 70B model on a single GPU — with 4.6 TB/s bandwidth and 1.8 PFLOPS FP8 performance. The GB200 NVL72 rack-scale system pairs 36 Grace CPUs with 72 B200 GPUs connected via NVLink Switch, delivering 1.4 exaFLOPS FP8 for training. Blackwell cloud availability is expanding in 2026 but remains limited compared to Hopper. B200 on-demand pricing is currently $8–$12/hr where available.
NVIDIA Ada Lovelace: L40S, L4, RTX 4090
Ada Lovelace (2022–2023) targets inference and professional workloads. The L40S (48GB GDDR6, 733 TFLOPS FP16) is the primary data-centre inference GPU for 7B–34B models, offering a better cost-per-token than H100 for models that fit in 48GB. The L4 (24GB GDDR6, 242 TFLOPS FP16) is the most cost-effective option for 7B model inference and video processing. The RTX 4090 (24GB GDDR6X, 330 TFLOPS FP16) is popular for fine-tuning and development — available from $0.44/hr on Vast.ai and RunPod, making it the cheapest path to 24GB VRAM in the cloud.
NVIDIA Ampere: A100, A40, A10, A10G
Ampere (2020–2021) remains widely available and cost-effective for many AI workloads. The A100 80GB SXM (312 TFLOPS FP16, 2 TB/s bandwidth, 600 GB/s NVLink 3.0) is the previous-generation training standard — still capable for 13B–70B model training at $1.89–$3.20/hr, roughly 40–60% cheaper than H100. The A100 40GB is available from $1.50/hr. The A40 (48GB GDDR6) and A10/A10G (24GB GDDR6) are popular inference GPUs — A10G is the standard GPU in AWS g5 instances. For workloads that don't require H100's bandwidth or NVLink scale, Ampere offers the best price-to-performance ratio in 2026.
AMD Instinct: MI300X and MI355X
AMD's Instinct MI300X (192GB HBM3, 5.3 TB/s bandwidth, 1.3 PFLOPS FP16) is the primary NVIDIA alternative for large-model inference. Its 192GB memory capacity — 2.4× the H100's 80GB — enables serving 70B models on a single GPU and 405B models on a 4-GPU node, reducing the multi-GPU overhead that H100 requires for large models. The MI355X (288GB HBM3e) extends this advantage further. ROCm software support has improved significantly in 2025–2026, with PyTorch, vLLM, and TGI all supporting AMD GPUs. Cloud availability is growing — Oracle Cloud, Fluidstack, and TensorWave offer MI300X instances.
Key specs to compare: VRAM and bandwidth
For AI workloads, VRAM capacity and memory bandwidth are often more important than raw FLOPS. VRAM determines the maximum model size that fits on a single GPU — a 70B model in FP16 requires 140GB, necessitating multi-GPU configurations on H100 (80GB) but fitting on a single MI300X (192GB). Memory bandwidth determines how fast weights can be loaded during inference — the primary bottleneck for autoregressive generation. H100 SXM's 3.35 TB/s bandwidth enables 2× faster inference than A100's 2 TB/s for the same model, even at similar FLOPS.
FP8 and quantisation in 2026
FP8 precision (supported on H100, H200, B200) doubles effective throughput versus FP16 for inference with minimal quality degradation. A 70B model in FP8 requires 70GB VRAM — fitting on a single H100 80GB — versus 140GB in FP16. INT4 quantisation (GPTQ, AWQ, GGUF) reduces this further to 35GB, enabling 70B inference on 2× A100 40GB or a single MI300X. The practical implication: for inference workloads, effective VRAM after quantisation is the relevant constraint, not raw VRAM capacity. The GPU specs database shows raw VRAM; factor in your target precision when selecting hardware.
NVLink and multi-GPU scaling
NVLink bandwidth determines how efficiently multiple GPUs can collaborate on a single large model. H100 SXM's NVLink 4.0 provides 900 GB/s bidirectional bandwidth between GPUs in an 8-GPU node — fast enough to treat 8× H100 as a single 640GB memory pool for tensor parallelism. H100 PCIe uses PCIe 5.0 (128 GB/s) instead of NVLink, making it significantly less efficient for multi-GPU training. A100 SXM uses NVLink 3.0 (600 GB/s). For single-GPU inference or training jobs that fit on one GPU, NVLink is irrelevant — PCIe variants are cheaper and equally capable.