H100 40GB
PCIe variant of H100. Lower TDP and cost than SXM5. Good for inference-heavy workloads.
H100 40GB Overview
The H100 40GB is a PCIe-form-factor Hopper GPU that brings the Transformer Engine and FP8 support to standard server deployments. It combines 40GB of HBM3 with 1,513 TFLOPS of FP16/BF16 performance and 3,026 TFLOPS at FP8, at a 350W power target that is easier to deploy than an SXM H100.
Its 2,000 GB/s HBM3 bandwidth is strong for a PCIe accelerator and materially improves transformer inference over older Ampere cards. The trade-off is capacity: 40GB limits native FP16 placement to roughly 20B parameters before runtime overhead, KV cache, and batch size are considered, even though the compute engine is much faster than that ceiling suggests.
This model fits mid-size LLM inference, FP8-aware serving, and training jobs that favor PCIe density over an SXM platform. It is not the best choice for single-card 70B-class FP16 inference or applications whose context cache consumes most of the 40GB allocation.
Memory
Compute Performance
Hardware
Relative Performance
Relative to highest-spec GPU in database
Limitations
Live Cloud PricingOn-demand hourly rates
Compare H100 40GB vs…
Use Case Guidance
LLM Model Size Guidance
Estimates only. Actual capacity depends on context length, KV cache, and framework overhead.
Related Guides
LLM APIs Running on This GPU Class
Providers that serve frontier LLM inference on Hopper-class hardware.
Related GPUs
Frequently Asked Questions
How much VRAM does the H100 40GB have?
The H100 40GB has 40GB of HBM3 memory with 2000 GB/s bandwidth. This enables running models up to approximately 80B parameters at INT4 precision, 40B at INT8, or 20B at FP16.
What is the FP16 performance of the H100 40GB?
The H100 40GB delivers 1513 TFLOPS of FP16 performance and 1513 TFLOPS BF16, and 3026 TFLOPS FP8. INT8 throughput is 3026 TOPS. For transformer inference, memory bandwidth (2000 GB/s) is often the binding constraint rather than raw TFLOPS.
What is the H100 40GB best used for?
The H100 40GB is best suited for: Mid-size LLM inference, Cost-efficient training, PCIe deployments. PCIe variant of H100. Lower TDP and cost than SXM5. Good for inference-heavy workloads.
What interconnect does the H100 40GB use?
The H100 40GB uses NVLink 4.0 / PCIe 5.0 with 900 GB/s NVLink bandwidth for multi-GPU configurations. NVLink enables near-linear tensor-parallel scaling across multiple cards for models that exceed single-card VRAM.
What LLM model sizes can the H100 40GB run?
With 40GB of HBM3, the H100 40GB can run models up to approximately 20B parameters at FP16 (2 bytes/param), 40B at INT8 (1 byte/param), or 80B at INT4/GGUF (0.5 bytes/param). These are estimates — actual capacity depends on context length, KV cache size, and framework overhead. Longer context windows require more KV cache memory, reducing the effective model size that fits.
How does the H100 40GB compare to the A100 for LLM inference?
The H100 40GB has 1513 TFLOPS FP16 vs the A100 80GB's 312 TFLOPS, and 2000 GB/s memory bandwidth vs the A100's 2,039 GB/s. For memory-bound autoregressive LLM inference, bandwidth is the primary determinant of tokens-per-second. The A100's higher bandwidth gives it a throughput advantage for large model inference, despite the H100 40GB's lower cost.
What is the power consumption of the H100 40GB?
The H100 40GB has a TDP (Thermal Design Power) of 350W. This is the maximum sustained power draw under full load. For data center deployments, total rack power consumption is typically 1.2–1.5× the GPU TDP when accounting for CPU, memory, networking, and cooling overhead. At 350W, the H100 40GB is in the mid-range tier — compatible with standard data center power infrastructure.
Ready to rent?
Compare H100 40GB prices across 97+ providers
Live on-demand & spot rates · monthly cost estimates · availability status