Rent GH200 Grace Hopper 96GB
Compare live on-demand and spot rental prices across 97+ cloud providers. Grace Hopper superchip — ARM Grace CPU + H100 GPU connected via NVLink-C2C at 900 GB/s. 96GB HBM3 + 480GB LPDDR5X system memory accessible as unified pool.
Choosing the right billing model for GH200 Grace Hopper 96GB
Provision and terminate at any time. Ideal for development, short experiments, and workloads with unpredictable duration.
Instances can be reclaimed when demand spikes. Best for fault-tolerant batch jobs, training with checkpointing, and preprocessing.
Lock in a rate for 1–3 months. Right for sustained production inference or long training runs where cost predictability matters.
Best use cases
- Memory-bound LLM inference
- Large context windows
- Unified CPU+GPU workloads
GH200 Grace Hopper 96GB — Specs & Benchmarks
Performance bars, compute tiers (FP32/FP16/BF16/FP8/INT8), memory specs, LLM model size guidance, and related GPU comparisons.
GH200 Grace Hopper 96GB Rental Guide
Rent GH200 for unified-memory workflows, very large contexts, or Hopper inference that benefits from the Grace CPU’s close coupling. It is a platform choice rather than simply a 96GB GPU rental, so validate software support, CPU architecture requirements, and memory-placement behavior before committing.
Spot capacity can be economical for checkpointed simulation, batch inference, and preprocessing jobs, but GH200 supply is specialized and less interchangeable than standard GPU instances. Choose on-demand or reserved systems for production services that depend on the Grace-Hopper topology and continuous access.
Compare GH200 with H200 for maximum HBM capacity and with standard H100 systems for broader availability. The Grace CPU and NVLink-C2C are valuable where unified memory reduces complexity; for compute-bound jobs that already fit on an H100, the premium and deployment constraints may not pay back.
Frequently Asked Questions
How much does it cost to rent a GH200 Grace Hopper 96GB?
GH200 Grace Hopper 96GB on-demand rental prices vary by provider and region. On-demand rates typically range based on availability and provider margins — use the comparison table above to see current live rates across all providers. Spot instances are generally 40–70% cheaper than on-demand but can be interrupted. Monthly cost estimates (hourly rate × 730 hours) are shown in the table for sustained workloads.
Which cloud provider has the cheapest GH200 Grace Hopper 96GB?
The cheapest GH200 Grace Hopper 96GB provider changes as providers update their pricing. The comparison table above shows live rates sorted by price, so the cheapest option is always at the top. Factors beyond headline price include region (latency to your users), availability (high/medium/low), and billing granularity (per-second vs per-hour minimums).
What can I run on a GH200 Grace Hopper 96GB?
With 96GB of HBM3, the GH200 Grace Hopper 96GB can run LLM models up to approximately 48B parameters at FP16, 96B at INT8, or 192B at INT4/GGUF quantization. Common workloads include: Memory-bound LLM inference, Large context windows, Unified CPU+GPU workloads. Grace Hopper superchip — ARM Grace CPU + H100 GPU connected via NVLink-C2C at 900 GB/s. 96GB HBM3 + 480GB LPDDR5X system memory accessible as unified pool.
Should I use on-demand or spot pricing for GH200 Grace Hopper 96GB?
Spot instances save 40–70% vs on-demand but can be interrupted when the provider needs capacity back. Use spot for: batch inference jobs, training runs with checkpointing, preprocessing pipelines, and any workload that can tolerate interruption and restart. Use on-demand for: production inference serving, interactive workloads, and jobs that cannot be interrupted. Most providers bill per second, so short on-demand jobs are not penalized by hourly minimums.
How does the GH200 Grace Hopper 96GB compare to the H100 for cloud rental?
The H100 80GB delivers 1,979 TFLOPS FP16 with 3,350 GB/s HBM3 bandwidth, compared to the GH200 Grace Hopper 96GB's 1979 TFLOPS FP16 and 4000 GB/s bandwidth. The H100 is significantly more expensive — typically $2.50–$5.00/hr vs lower rates for the GH200 Grace Hopper 96GB. For workloads that fit within 96GB and don't require FP8 precision, the GH200 Grace Hopper 96GB often delivers better cost-per-token than the H100.
What is the memory bandwidth of the GH200 Grace Hopper 96GB and why does it matter?
The GH200 Grace Hopper 96GB has 4000 GB/s of memory bandwidth. For LLM inference, memory bandwidth is often more important than raw TFLOPS — each autoregressive token generation reads the full model weight matrix from VRAM, so bandwidth directly determines tokens-per-second throughput. Higher bandwidth means faster inference for the same model at the same batch size. For batch inference (processing many requests simultaneously), compute throughput becomes more important.
Can I use the GH200 Grace Hopper 96GB for Stable Diffusion or image generation?
Yes — the GH200 Grace Hopper 96GB is well-suited for Stable Diffusion and image generation workloads. Image generation is primarily FP32 and FP16 compute-bound, and the GH200 Grace Hopper 96GB's 67 TFLOPS FP32 throughput determines images-per-second. The 96GB VRAM fits SDXL (requires ~6GB) and most ControlNet pipelines. For high-throughput image generation at scale, compare cost-per-image across providers using the GPU cost calculator.