Compute Comparison
NVIDIAHopper2023

H100 40GB

PCIe variant of H100. Lower TDP and cost than SXM5. Good for inference-heavy workloads.

VRAM
40GB
HBM3
FP16
1.5k
TFLOPS
Bandwidth
2.0k
GB/s
TDP
350W
power
Best for:Mid-size LLM inferenceCost-efficient trainingPCIe deployments

H100 40GB Overview

The H100 40GB is a PCIe-form-factor Hopper GPU that brings the Transformer Engine and FP8 support to standard server deployments. It combines 40GB of HBM3 with 1,513 TFLOPS of FP16/BF16 performance and 3,026 TFLOPS at FP8, at a 350W power target that is easier to deploy than an SXM H100.

Its 2,000 GB/s HBM3 bandwidth is strong for a PCIe accelerator and materially improves transformer inference over older Ampere cards. The trade-off is capacity: 40GB limits native FP16 placement to roughly 20B parameters before runtime overhead, KV cache, and batch size are considered, even though the compute engine is much faster than that ceiling suggests.

This model fits mid-size LLM inference, FP8-aware serving, and training jobs that favor PCIe density over an SXM platform. It is not the best choice for single-card 70B-class FP16 inference or applications whose context cache consumes most of the 40GB allocation.

Memory

VRAM40 GB
Memory TypeHBM3
Bandwidth2000 GB/s
NVLink BW900 GB/s

Compute Performance

FP3251 TFLOPS
FP161513 TFLOPS
BF161513 TFLOPS
FP83026 TFLOPS
INT83026 TOPS

Hardware

ArchitectureGH100
GenerationHopper
Process NodeTSMC 4N
Transistors80B
TDP350 W
InterconnectNVLink 4.0 / PCIe 5.0
Release Year2023

Relative Performance

FP16 Compute20%
VRAM Capacity14%
Mem Bandwidth13%

Relative to highest-spec GPU in database

Limitations

Only 40GB VRAM — limits to ~30B models at FP16
Lower bandwidth than SXM5 variant (2 TB/s vs 3.35 TB/s)
Higher cost than A100 for inference-only workloads

Live Cloud PricingOn-demand hourly rates

Loading live prices…

Compare H100 40GB vs…

Use Case Guidance

Mid-size LLM inference
Cost-efficient training
PCIe deployments

LLM Model Size Guidance

Max model (FP16)~20Bparameters at FP16 precision
Max model (INT8)~40Bparameters at INT8 precision
Max model (INT4)~80Bparameters at INT4/GGUF

Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.

Related Guides

LLM APIs Running on This GPU Class

Providers that serve frontier LLM inference on Hopper-class hardware.

Browse all 42 LLM models

Related GPUs

Frequently Asked Questions

How much VRAM does the H100 40GB have?

The H100 40GB has 40GB of HBM3 memory with 2000 GB/s bandwidth. This enables running models up to approximately 80B parameters at INT4 precision, 40B at INT8, or 20B at FP16.

What is the FP16 performance of the H100 40GB?

The H100 40GB delivers 1513 TFLOPS of FP16 performance and 1513 TFLOPS BF16, and 3026 TFLOPS FP8. INT8 throughput is 3026 TOPS. For transformer inference, memory bandwidth (2000 GB/s) is often the binding constraint rather than raw TFLOPS.

What is the H100 40GB best used for?

The H100 40GB is best suited for: Mid-size LLM inference, Cost-efficient training, PCIe deployments. PCIe variant of H100. Lower TDP and cost than SXM5. Good for inference-heavy workloads.

What interconnect does the H100 40GB use?

The H100 40GB uses NVLink 4.0 / PCIe 5.0 with 900 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.

What LLM model sizes can the H100 40GB run?

With 40GB of HBM3, the H100 40GB can run models up to approximately 20B parameters at FP16 (2 bytes/param), 40B at INT8 (1 byte/param), or 80B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.

How does the H100 40GB compare to the A100 for LLM inference?

The H100 40GB has 1513 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 2000 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the H100 40GB's lower cost.

What is the power consumption of the H100 40GB?

The H100 40GB has a TDP (Thermal Design Power) of 350W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 350W, the H100 40GB is in the mid-range tier — compatible with standard data center power infrastructure.

Ready to rent?

Compare H100 40GB prices across 97+ providers

Live on-demand & spot rates · monthly cost estimates · availability status

Compare rental prices