Rent H100 80GB
Compare live on-demand and spot rental prices across 97+ cloud providers. Flagship data center GPU. Transformer Engine with FP8 support. SXM5 form factor for max bandwidth.
Choosing the right billing model for H100 80GB
Provision and terminate at any time. Ideal for development, short experiments, and workloads with unpredictable duration.
Instances can be reclaimed when demand spikes. Best for fault-tolerant batch jobs, training with checkpointing, and preprocessing.
Lock in a rate for 1–3 months. Right for sustained production inference or long training runs where cost predictability matters.
Best use cases
- LLM training
- Large-scale inference
- Scientific HPC
H100 80GB — Specs & Benchmarks
Performance bars, compute tiers (FP32/FP16/BF16/FP8/INT8), memory specs, LLM model size guidance, and related GPU comparisons.
H100 80GB Rental Guide
Renting an H100 80GB is justified when you need to train models above 30B parameters, run 70B+ inference at scale, or use FP8 quantization to maximize throughput. The 80GB HBM3 fits a 70B model at FP16 (140GB requires two H100s with tensor parallelism) or a 140B model at INT4. For inference serving of models like Llama 3.1 70B or Mixtral 8×22B, the H100's 3,350 GB/s memory bandwidth delivers roughly 2× the tokens-per-second of an A100 80GB at similar cost-per-hour.
Spot instances for H100 are available from providers including Lambda Labs, CoreWeave, and Vast.ai at 40–65% below on-demand rates. Spot is well-suited for training runs with checkpointing, batch inference jobs, and preprocessing pipelines. For interactive inference serving where uptime matters, on-demand or reserved instances are more appropriate. Most providers bill per second, so short benchmark runs are not penalized.
When comparing H100 providers, check whether the listing is SXM5 (900 GB/s NVLink, 3,350 GB/s HBM3) or PCIe (the H100 PCIe has 2,000 GB/s bandwidth and lower NVLink bandwidth). For multi-GPU tensor parallelism, SXM5 is significantly faster. For single-card inference, the PCIe variant is often 20–30% cheaper with minimal throughput difference. Use the GPU cost calculator to model total cost across your expected utilization before committing to a provider.
Frequently Asked Questions
How much does it cost to rent a H100 80GB?
H100 80GB on-demand rental prices vary by provider and region. On-demand rates typically range based on availability and provider margins — use the comparison table above to see current live rates across all providers. Spot instances are generally 40–70% cheaper than on-demand but can be interrupted. Monthly cost estimates (hourly rate × 730 hours) are shown in the table for sustained workloads.
Which cloud provider has the cheapest H100 80GB?
The cheapest H100 80GB provider changes as providers update their pricing. The comparison table above shows live rates sorted by price, so the cheapest option is always at the top. Factors beyond headline price include region (latency to your users), availability (high/medium/low), and billing granularity (per-second vs per-hour minimums).
What can I run on a H100 80GB?
With 80GB of HBM3, the H100 80GB can run LLM models up to approximately 40B parameters at FP16, 80B at INT8, or 160B at INT4/GGUF quantization. Common workloads include: LLM training, Large-scale inference, Scientific HPC. Flagship data center GPU. Transformer Engine with FP8 support. SXM5 form factor for max bandwidth.
Should I use on-demand or spot pricing for H100 80GB?
Spot instances save 40–70% vs on-demand but can be interrupted when the provider needs capacity back. Use spot for: batch inference jobs, training runs with checkpointing, preprocessing pipelines, and any workload that can tolerate interruption and restart. Use on-demand for: production inference serving, interactive workloads, and jobs that cannot be interrupted. Most providers bill per second, so short on-demand jobs are not penalized by hourly minimums.
How does the H100 80GB compare to the H100 for cloud rental?
The H100 80GB delivers 1,979 TFLOPS FP16 with 3,350 GB/s HBM3 bandwidth, compared to the H100 80GB's 1979 TFLOPS FP16 and 3350 GB/s bandwidth. The H100 is significantly more expensive — typically $2.50–$5.00/hr vs lower rates for the H100 80GB. For workloads that fit within 80GB and don't require FP8 precision, the H100 80GB often delivers better cost-per-token than the H100.
What is the memory bandwidth of the H100 80GB and why does it matter?
The H100 80GB has 3350 GB/s of memory bandwidth. For LLM inference, memory bandwidth is often more important than raw TFLOPS — each autoregressive token generation reads the full model weight matrix from VRAM, so bandwidth directly determines tokens-per-second throughput. Higher bandwidth means faster inference for the same model at the same batch size. For batch inference (processing many requests simultaneously), compute throughput becomes more important.
Can I use the H100 80GB for Stable Diffusion or image generation?
Yes — the H100 80GB is well-suited for Stable Diffusion and image generation workloads. Image generation is primarily FP32 and FP16 compute-bound, and the H100 80GB's 67 TFLOPS FP32 throughput determines images-per-second. The 80GB VRAM fits SDXL (requires ~6GB) and most ControlNet pipelines. For high-throughput image generation at scale, compare cost-per-image across providers using the GPU cost calculator.