Compute Comparison
87 GPUsUpdated July 2026

GPU Specs Database — H100, A100, RTX 4090 & All AI Cloud GPU Specs

Complete technical specifications for every major AI and cloud GPU — NVIDIA Hopper, Blackwell, Ampere, Ada Lovelace, AMD CDNA 3, and more. Click any GPU to see live cloud pricing across 94+ providers.

Total GPUs
87
Generations
14
Max FP16
7.5kT
Max VRAM
288GB
87 results
Sort by:
NVIDIABlackwell Ultra2026

GB300 NVL72

FP16 Performance7.5k TFLOPS
VRAM288GB HBM3e
BW16000 GB/s
FP32120T
TDP1200W
Next-gen frontier trainingTrillion-param inference+1
NVIDIABlackwell2025

GB200 NVL72

FP16 Performance4.5k TFLOPS
VRAM192GB HBM3e
BW8000 GB/s
FP3290T
TDP1000W
Frontier LLM trainingTrillion-param models+1
NVIDIABlackwell2025

B200 192GB

FP16 Performance3.5k TFLOPS
VRAM192GB HBM3e
BW8000 GB/s
FP3280T
TDP1000W
Next-gen LLM trainingUltra-large model inference+1
NVIDIABlackwell2025

B200 SXM 192GB

FP16 Performance3.5k TFLOPS
VRAM192GB HBM3e
BW8000 GB/s
FP3280T
TDP1000W
Next-gen LLM trainingUltra-large model inference+1
NVIDIABlackwell2025

B200 NVL 192GB

FP16 Performance3.5k TFLOPS
VRAM192GB HBM3e
BW8000 GB/s
FP3280T
TDP900W
Large-scale inferenceMulti-GPU NVLink clusters+1
NVIDIABlackwell Ultra2025

B300 SXM6

FP16 Performance2.8k TFLOPS
VRAM288GB HBM3e
BW8000 GB/s
FP3260T
TDP1000W
Frontier LLM trainingMassive multi-GPU clusters+1
NVIDIABlackwell Ultra2025

GB300 SXM6

FP16 Performance2.8k TFLOPS
VRAM288GB HBM3e
BW8000 GB/s
FP3260T
TDP1200W
Frontier AI trainingInference at scale+1
NVIDIABlackwell Ultra2025

B300 SXM 262GB

FP16 Performance2.8k TFLOPS
VRAM262GB HBM3e
BW8000 GB/s
FP3260T
TDP1000W
Frontier LLM trainingMassive multi-GPU clusters+1
NVIDIABlackwell Ultra2025

GB300 SXM 288GB

FP16 Performance2.8k TFLOPS
VRAM288GB HBM3e
BW8000 GB/s
FP3260T
TDP1200W
Frontier AI trainingInference at scale+1
NVIDIAHopper2024

H200 141GB

FP16 Performance2.0k TFLOPS
VRAM141GB HBM3e
BW4800 GB/s
FP3267T
TDP700W
70B+ model inferenceMemory-bound LLM serving+1
NVIDIAHopper2022

H100 80GB

FP16 Performance2.0k TFLOPS
VRAM80GB HBM3
BW3350 GB/s
FP3267T
TDP700W
LLM trainingLarge-scale inference+1
NVIDIAHopper2022

H100 SXM5 80GB

FP16 Performance2.0k TFLOPS
VRAM80GB HBM3
BW3350 GB/s
FP3267T
TDP700W
LLM trainingLarge-scale inference+1
NVIDIAHopper2023

H100 NVL 94GB

FP16 Performance2.0k TFLOPS
VRAM94GB HBM3
BW3938 GB/s
FP3267T
TDP400W
Large context LLM inferenceMemory-bound workloads+1
NVIDIAHopper2024

H200 SXM 141GB

FP16 Performance2.0k TFLOPS
VRAM141GB HBM3e
BW4800 GB/s
FP3267T
TDP700W
70B+ model inferenceMemory-bound LLM serving+1
NVIDIAHopper2024

H200 NVL 141GB

FP16 Performance2.0k TFLOPS
VRAM141GB HBM3e
BW4800 GB/s
FP3267T
TDP600W
Large model inferenceMulti-GPU NVLink clusters+1
NVIDIAHopper2023

GH200 Grace Hopper 96GB

FP16 Performance2.0k TFLOPS
VRAM96GB HBM3
BW4000 GB/s
FP3267T
TDP900W
Memory-bound LLM inferenceLarge context windows+1
NVIDIAHopper2023

H100 40GB

FP16 Performance1.5k TFLOPS
VRAM40GB HBM3
BW2000 GB/s
FP3251T
TDP350W
Mid-size LLM inferenceCost-efficient training+1
NVIDIAHopper2022

H100 PCIe 80GB

FP16 Performance1.5k TFLOPS
VRAM80GB HBM3
BW2000 GB/s
FP3251T
TDP350W
Inference servingCost-efficient training+1
AMDCDNA 3.52024

MI325X 256GB

FP16 Performance1.5k TFLOPS
VRAM256GB HBM3e
BW6000 GB/s
FP32183.6T
TDP750W
Massive context LLM inferenceMemory-bound workloads+1
AMDCDNA 3.52024

MI325X 256GB

FP16 Performance1.5k TFLOPS
VRAM256GB HBM3e
BW6000 GB/s
FP32183.6T
TDP750W
Massive context LLM inferenceMemory-bound workloads+1
AMDCDNA 32023

MI300X 192GB

FP16 Performance1.3k TFLOPS
VRAM192GB HBM3
BW5300 GB/s
FP32163.4T
TDP750W
Memory-bound LLM inferenceLarge context windows+1
AMDCDNA 32023

MI300X 192GB

FP16 Performance1.3k TFLOPS
VRAM192GB HBM3
BW5300 GB/s
FP32163.4T
TDP750W
Memory-bound LLM inferenceLarge context windows+1
AMDCDNA 42025

MI355X

FP16 Performance1.3k TFLOPS
VRAM288GB HBM3e
BW8000 GB/s
FP32163.4T
TDP750W
Large-scale LLM trainingMemory-bound inference+1
AMDCDNA 42025

MI355X 288GB

FP16 Performance1.3k TFLOPS
VRAM288GB HBM3e
BW8000 GB/s
FP32163.4T
TDP750W
Large-scale LLM trainingMemory-bound inference+1
NVIDIAGaudi 22022

Gaudi 2 96GB

FP16 Performance432 TFLOPS
VRAM96GB HBM2e
BW2457 GB/s
FP327T
TDP600W
Cost-effective LLM trainingTransformer workloads+1
NVIDIAGaudi 22022

Gaudi 2

FP16 Performance432 TFLOPS
VRAM96GB HBM2e
BW2457 GB/s
FP327T
TDP600W
Cost-effective LLM trainingTransformer workloads+1
NVIDIABlackwell2025

RTX 5090

FP16 Performance419.6 TFLOPS
VRAM32GB GDDR7
BW1792 GB/s
FP32209.8T
TDP575W
Consumer AI inferenceFine-tuning up to 30B+1
AMDCDNA 22021

MI250X 128GB

FP16 Performance383 TFLOPS
VRAM128GB HBM2e
BW3277 GB/s
FP3247.9T
TDP560W
Large model trainingMemory-bound HPC+1
NVIDIAAmpere2020

A100 80GB

FP16 Performance312 TFLOPS
VRAM80GB HBM2e
BW2000 GB/s
FP3219.5T
TDP400W
ML trainingLarge model inference+1
NVIDIAAmpere2020

A100 40GB

FP16 Performance312 TFLOPS
VRAM40GB HBM2e
BW1555 GB/s
FP3219.5T
TDP300W
ML trainingMid-size model inference+1
NVIDIAAmpere2020

A100 SXM4 80GB

FP16 Performance312 TFLOPS
VRAM80GB HBM2e
BW2000 GB/s
FP3219.5T
TDP400W
ML trainingLarge model inference+1
NVIDIAAmpere2021

A100 PCIe 80GB

FP16 Performance312 TFLOPS
VRAM80GB HBM2e
BW1935 GB/s
FP3219.5T
TDP300W
Inference servingCost-efficient training+1
NVIDIAAmpere2020

A100 PCIe 40GB

FP16 Performance312 TFLOPS
VRAM40GB HBM2e
BW1555 GB/s
FP3219.5T
TDP250W
Mid-size model inferenceCost-efficient training+1
NVIDIAAmpere2020

A100 SXM4 40GB

FP16 Performance312 TFLOPS
VRAM40GB HBM2e
BW1555 GB/s
FP3219.5T
TDP300W
ML trainingMulti-GPU NVLink setups+1
NVIDIABlackwell2025

RTX 5080

FP16 Performance275.4 TFLOPS
VRAM16GB GDDR7
BW960 GB/s
FP32137.7T
TDP360W
Budget consumer inferenceSmall model fine-tuning+1
NVIDIABlackwell2025

RTX PRO 6000 Blackwell

FP16 Performance250 TFLOPS
VRAM96GB GDDR7
BW1792 GB/s
FP32125T
TDP300W
Professional AI workloadsLarge model inference+1
NVIDIABlackwell2025

RTX 5070 Ti

FP16 Performance215.2 TFLOPS
VRAM16GB GDDR7
BW896 GB/s
FP32107.6T
TDP300W
Budget inferenceExperimentation+1
NVIDIAAda Lovelace2023

L40S

FP16 Performance183 TFLOPS
VRAM48GB GDDR6
BW864 GB/s
FP3291.6T
TDP350W
Inference servingGenerative AI+1
NVIDIAAda Lovelace2022

RTX 6000 Ada

FP16 Performance182.2 TFLOPS
VRAM48GB GDDR6
BW960 GB/s
FP3291.1T
TDP300W
Large model inferenceMulti-GPU professional workloads+1
NVIDIAAda Lovelace2022

L40

FP16 Performance181 TFLOPS
VRAM48GB GDDR6
BW864 GB/s
FP3290.5T
TDP300W
VisualizationInference+1
NVIDIAAda Lovelace2024

RTX 5880 Ada

FP16 Performance181 TFLOPS
VRAM48GB GDDR6
BW960 GB/s
FP3290.5T
TDP285W
Professional inferenceLarge model workstation use+1
NVIDIAAda Lovelace2022

RTX 4090

FP16 Performance165.2 TFLOPS
VRAM24GB GDDR6X
BW1008 GB/s
FP3282.6T
TDP450W
Fine-tuning small modelsInference up to 13B+1
NVIDIAAmpere2021

A30

FP16 Performance165 TFLOPS
VRAM24GB HBM2
BW933 GB/s
FP3210.3T
TDP165W
Inference at scaleLow-power HPC+1
NVIDIABlackwell2025

RTX 5070

FP16 Performance144.2 TFLOPS
VRAM12GB GDDR7
BW672 GB/s
FP3272.1T
TDP250W
Ultra-budget inferenceExperimentation+1
NVIDIAAda Lovelace2023

RTX 5000 Ada

FP16 Performance130.6 TFLOPS
VRAM32GB GDDR6
BW576 GB/s
FP3265.3T
TDP250W
Mid-range professional inference30B model serving+1
NVIDIATuring2018

Quadro RTX 6000 24GB

FP16 Performance130.5 TFLOPS
VRAM24GB GDDR6
BW672 GB/s
FP3216.3T
TDP295W
Budget 13B inferenceNVLink multi-GPU setups+1
NVIDIATuring2018

Quadro RTX 8000 48GB

FP16 Performance130.5 TFLOPS
VRAM48GB GDDR6
BW672 GB/s
FP3216.3T
TDP295W
Budget 30B inferenceLarge VRAM workloads+1
NVIDIAAmpere2021

A10G

FP16 Performance125 TFLOPS
VRAM24GB GDDR6
BW600 GB/s
FP3231.2T
TDP150W
Inference servingGraphics workloads+1
NVIDIAAmpere2021

A10 24GB

FP16 Performance125 TFLOPS
VRAM24GB GDDR6
BW600 GB/s
FP3231.2T
TDP150W
Inference servingGraphics + compute+1
NVIDIAVolta2018

V100 32GB

FP16 Performance112 TFLOPS
VRAM32GB HBM2
BW900 GB/s
FP3214T
TDP300W
Legacy ML trainingHPC workloads+1
NVIDIAVolta2017

V100 16GB

FP16 Performance112 TFLOPS
VRAM16GB HBM2
BW900 GB/s
FP3214T
TDP250W
Legacy ML trainingBudget inference+1
NVIDIAVolta2017

TITAN V 12GB

FP16 Performance110 TFLOPS
VRAM12GB HBM2
BW653 GB/s
FP3213.8T
TDP250W
Legacy ML researchBudget Volta workloads+1
NVIDIAAda Lovelace2022

RTX 4080

FP16 Performance97.5 TFLOPS
VRAM16GB GDDR6X
BW717 GB/s
FP3248.7T
TDP320W
Budget inferenceSmall model fine-tuning+1
NVIDIAAda Lovelace2023

RTX 4500 Ada

FP16 Performance97.4 TFLOPS
VRAM24GB GDDR6
BW432 GB/s
FP3248.7T
TDP210W
Enterprise inferenceBudget professional AI+1
NVIDIATuring2018

Quadro RTX 5000 16GB

FP16 Performance89.2 TFLOPS
VRAM16GB GDDR6
BW448 GB/s
FP3211.2T
TDP265W
Budget inferenceNVLink multi-GPU setups+1
NVIDIAAda Lovelace2023

RTX 4070 Ti

FP16 Performance80.2 TFLOPS
VRAM12GB GDDR6X
BW504 GB/s
FP3240.1T
TDP285W
Budget inferenceSmall model fine-tuning+1
NVIDIAAmpere2020

RTX A6000

FP16 Performance77.4 TFLOPS
VRAM48GB GDDR6
BW768 GB/s
FP3238.7T
TDP300W
Large model inferenceMulti-GPU NVLink setups+1
NVIDIAAmpere2020

A40

FP16 Performance74.8 TFLOPS
VRAM48GB GDDR6
BW696 GB/s
FP3237.4T
TDP300W
Visualization + computeMid-size inference+1
NVIDIAAmpere2020

RTX 3090

FP16 Performance71 TFLOPS
VRAM24GB GDDR6X
BW936 GB/s
FP3235.6T
TDP350W
Budget inferenceSmall model fine-tuning+1
NVIDIABlackwell2025

RTX 5060 Ti

FP16 Performance70.4 TFLOPS
VRAM16GB GDDR7
BW672 GB/s
FP3235.2T
TDP180W
Budget AI inferenceSmall model serving+1
NVIDIAAmpere2021

RTX 3080 Ti

FP16 Performance68.2 TFLOPS
VRAM12GB GDDR6X
BW912 GB/s
FP3234.1T
TDP350W
Budget inferenceSmall model fine-tuning+1
NVIDIAAmpere2021

RTX A5500 24GB

FP16 Performance68.2 TFLOPS
VRAM24GB GDDR6
BW768 GB/s
FP3234.1T
TDP230W
Professional inferenceNVLink multi-GPU setups+1
NVIDIATuring2018

T4

FP16 Performance65 TFLOPS
VRAM16GB GDDR6
BW320 GB/s
FP328.1T
TDP70W
Cost-efficient inferenceHigh-density serving+1
NVIDIAAmpere2021

A16

FP16 Performance62.4 TFLOPS
VRAM64GB GDDR6
BW432 GB/s
FP3231.2T
TDP250W
VDI / virtual workstationsMulti-tenant inference+1
NVIDIAAda Lovelace2023

L4

FP16 Performance60.6 TFLOPS
VRAM24GB GDDR6
BW300 GB/s
FP3230.3T
TDP72W
Efficient inferenceEdge deployments+1
NVIDIAAmpere2020

RTX 3080

FP16 Performance59.6 TFLOPS
VRAM10GB GDDR6X
BW760 GB/s
FP3229.8T
TDP320W
Ultra-budget inferenceExperimentation+1
NVIDIAAda Lovelace2023

RTX 4070

FP16 Performance58.2 TFLOPS
VRAM12GB GDDR6X
BW504 GB/s
FP3229.1T
TDP200W
Ultra-budget inferenceExperimentation+1
NVIDIATuring2018

Quadro RTX 4000 8GB

FP16 Performance57 TFLOPS
VRAM8GB GDDR6
BW416 GB/s
FP327.1T
TDP160W
Ultra-budget inferenceSmall model experimentation+1
NVIDIAAmpere2021

RTX A5000

FP16 Performance55.6 TFLOPS
VRAM24GB GDDR6
BW768 GB/s
FP3227.8T
TDP230W
Professional inferenceBudget NVLink multi-GPU+1
NVIDIAAda Lovelace2023

RTX 4000 Ada

FP16 Performance53.4 TFLOPS
VRAM20GB GDDR6
BW360 GB/s
FP3226.7T
TDP130W
Entry enterprise inferenceLow-power deployments+1
NVIDIABlackwell2025

RTX 5060

FP16 Performance48 TFLOPS
VRAM8GB GDDR7
BW448 GB/s
FP3224T
TDP150W
Entry-level AI inferenceStable Diffusion+1
NVIDIAAmpere2021

RTX A4500 20GB

FP16 Performance47.4 TFLOPS
VRAM20GB GDDR6
BW640 GB/s
FP3223.7T
TDP200W
Professional inferenceMid-range AI workloads+1
NVIDIAAda Lovelace2023

RTX 4060 Ti 16GB

FP16 Performance44.2 TFLOPS
VRAM16GB GDDR6
BW288 GB/s
FP3222.1T
TDP165W
Budget inferenceSmall model fine-tuning+1
NVIDIAAmpere2020

RTX 3070 8GB

FP16 Performance40.6 TFLOPS
VRAM8GB GDDR6
BW448 GB/s
FP3220.3T
TDP220W
Ultra-budget inferenceSmall model experimentation+1
NVIDIAAda Lovelace2023

RTX 4000 SFF Ada

FP16 Performance38.4 TFLOPS
VRAM20GB GDDR6
BW272 GB/s
FP3219.2T
TDP70W
Ultra-low-power inferenceEdge deployments+1
NVIDIAAmpere2021

RTX A4000 16GB

FP16 Performance38.4 TFLOPS
VRAM16GB GDDR6
BW448 GB/s
FP3219.2T
TDP140W
Entry professional inferenceBudget AI workloads+1
NVIDIAAmpere2020

RTX 3060 Ti 8GB

FP16 Performance32.4 TFLOPS
VRAM8GB GDDR6
BW448 GB/s
FP3216.2T
TDP200W
Ultra-budget inferenceSmall model experimentation+1
NVIDIAAda Lovelace2023

RTX 4060 8GB

FP16 Performance30.2 TFLOPS
VRAM8GB GDDR6
BW272 GB/s
FP3215.1T
TDP115W
Ultra-budget inferenceSmall model experimentation+1
NVIDIAAmpere2021

RTX 3060 12GB

FP16 Performance25.4 TFLOPS
VRAM12GB GDDR6
BW360 GB/s
FP3212.7T
TDP170W
Ultra-budget inferenceSmall model experimentation+1
NVIDIAAda Lovelace2023

RTX 2000 Ada

FP16 Performance24 TFLOPS
VRAM16GB GDDR6
BW224 GB/s
FP3212T
TDP70W
Entry professional inferenceLow-power AI workloads+1
NVIDIAPascal2016

Tesla P100 16GB

FP16 Performance18.7 TFLOPS
VRAM16GB HBM2
BW732 GB/s
FP329.3T
TDP250W
Legacy ML workloadsBudget HPC+1
NVIDIAAmpere2021

RTX A2000 12GB

FP16 Performance16 TFLOPS
VRAM12GB GDDR6
BW288 GB/s
FP327.99T
TDP70W
Ultra-budget inferenceEdge AI+1
NVIDIAPascal2017

TITAN Xp 12GB

FP16 Performance12.1 TFLOPS
VRAM12GB GDDR5X
BW548 GB/s
FP3212.1T
TDP250W
Legacy inferenceUltra-budget workloads+1
NVIDIAPascal2016

Quadro P6000 24GB

FP16 Performance12 TFLOPS
VRAM24GB GDDR5X
BW432 GB/s
FP3212T
TDP250W
Legacy inferenceBudget 13B model serving+1
NVIDIAPascal2016

Quadro P5000 16GB

FP16 Performance8.9 TFLOPS
VRAM16GB GDDR5X
BW288 GB/s
FP328.9T
TDP180W
Legacy inferenceBudget 7B model serving+1
NVIDIAPascal2017

Quadro P4000 8GB

FP16 Performance5.3 TFLOPS
VRAM8GB GDDR5
BW243 GB/s
FP325.3T
TDP105W
Legacy inferenceUltra-budget workloads+1
NVIDIAMaxwell2015

Quadro M4000 8GB

FP16 Performance2.6 TFLOPS
VRAM8GB GDDR5
BW192 GB/s
FP322.6T
TDP120W
Legacy visualizationUltra-budget compute+1

GPU specs guide for AI and cloud compute

NVIDIA Hopper: H100 and H200

NVIDIA's Hopper architecture (2022–2023) remains the dominant GPU for large-scale AI training and inference in 2026. The H100 SXM delivers 989 TFLOPS FP16 with 80GB HBM3 at 3.35 TB/s bandwidth and 900 GB/s NVLink 4.0 interconnect — the standard for 70B+ LLM training. The H100 PCIe variant offers the same compute at lower bandwidth (2 TB/s) and is 10–20% cheaper on cloud. The H200 upgrades memory to 141GB HBM3e at 4.8 TB/s bandwidth — a significant advantage for inference serving of large models that previously required multi-GPU configurations. H100 on-demand pricing ranges from $3.50–$5.50/hr depending on provider.

NVIDIA Blackwell: B200 and GB200

NVIDIA's Blackwell architecture (2024–2025) delivers a 2.5–4× performance improvement over Hopper for AI inference. The B200 provides 180GB HBM3e memory — enough to serve a 70B model on a single GPU — with 4.6 TB/s bandwidth and 1.8 PFLOPS FP8 performance. The GB200 NVL72 rack-scale system pairs 36 Grace CPUs with 72 B200 GPUs connected via NVLink Switch, delivering 1.4 exaFLOPS FP8 for training. Blackwell cloud availability is expanding in 2026 but remains limited compared to Hopper. B200 on-demand pricing is currently $8–$12/hr where available.

NVIDIA Ada Lovelace: L40S, L4, RTX 4090

Ada Lovelace (2022–2023) targets inference and professional workloads. The L40S (48GB GDDR6, 733 TFLOPS FP16) is the primary data-centre inference GPU for 7B–34B models, offering a better cost-per-token than H100 for models that fit in 48GB. The L4 (24GB GDDR6, 242 TFLOPS FP16) is the most cost-effective option for 7B model inference and video processing. The RTX 4090 (24GB GDDR6X, 330 TFLOPS FP16) is popular for fine-tuning and development — available from $0.44/hr on Vast.ai and RunPod, making it the cheapest path to 24GB VRAM in the cloud.

NVIDIA Ampere: A100, A40, A10, A10G

Ampere (2020–2021) remains widely available and cost-effective for many AI workloads. The A100 80GB SXM (312 TFLOPS FP16, 2 TB/s bandwidth, 600 GB/s NVLink 3.0) is the previous-generation training standard — still capable for 13B–70B model training at $1.89–$3.20/hr, roughly 40–60% cheaper than H100. The A100 40GB is available from $1.50/hr. The A40 (48GB GDDR6) and A10/A10G (24GB GDDR6) are popular inference GPUs — A10G is the standard GPU in AWS g5 instances. For workloads that don't require H100's bandwidth or NVLink scale, Ampere offers the best price-to-performance ratio in 2026.

AMD Instinct: MI300X and MI355X

AMD's Instinct MI300X (192GB HBM3, 5.3 TB/s bandwidth, 1.3 PFLOPS FP16) is the primary NVIDIA alternative for large-model inference. Its 192GB memory capacity — 2.4× the H100's 80GB — enables serving 70B models on a single GPU and 405B models on a 4-GPU node, reducing the multi-GPU overhead that H100 requires for large models. The MI355X (288GB HBM3e) extends this advantage further. ROCm software support has improved significantly in 2025–2026, with PyTorch, vLLM, and TGI all supporting AMD GPUs. Cloud availability is growing — Oracle Cloud, Fluidstack, and TensorWave offer MI300X instances.

Key specs to compare: VRAM and bandwidth

For AI workloads, VRAM capacity and memory bandwidth are often more important than raw FLOPS. VRAM determines the maximum model size that fits on a single GPU — a 70B model in FP16 requires 140GB, necessitating multi-GPU configurations on H100 (80GB) but fitting on a single MI300X (192GB). Memory bandwidth determines how fast weights can be loaded during inference — the primary bottleneck for autoregressive generation. H100 SXM's 3.35 TB/s bandwidth enables 2× faster inference than A100's 2 TB/s for the same model, even at similar FLOPS.

FP8 and quantisation in 2026

FP8 precision (supported on H100, H200, B200) doubles effective throughput versus FP16 for inference with minimal quality degradation. A 70B model in FP8 requires 70GB VRAM — fitting on a single H100 80GB — versus 140GB in FP16. INT4 quantisation (GPTQ, AWQ, GGUF) reduces this further to 35GB, enabling 70B inference on 2× A100 40GB or a single MI300X. The practical implication: for inference workloads, effective VRAM after quantisation is the relevant constraint, not raw VRAM capacity. The GPU specs database shows raw VRAM; factor in your target precision when selecting hardware.

NVLink and multi-GPU scaling

NVLink bandwidth determines how efficiently multiple GPUs can collaborate on a single large model. H100 SXM's NVLink 4.0 provides 900 GB/s bidirectional bandwidth between GPUs in an 8-GPU node — fast enough to treat 8× H100 as a single 640GB memory pool for tensor parallelism. H100 PCIe uses PCIe 5.0 (128 GB/s) instead of NVLink, making it significantly less efficient for multi-GPU training. A100 SXM uses NVLink 3.0 (600 GB/s). For single-GPU inference or training jobs that fit on one GPU, NVLink is irrelevant — PCIe variants are cheaper and equally capable.