Compare GPU Cloud Providers — Side-by-Side H100, A100 & RTX Pricing Across 94+ Providers
Select any two providers to see live pricing, specs, regions, and strengths side-by-side.
Popular comparisons
All providers (97)
Click any provider to add it to the comparison. Click Compare to see the full side-by-side breakdown.
How to compare GPU cloud providers
What the comparison covers
The side-by-side comparison shows live on-demand and spot GPU pricing for every GPU model both providers offer, region availability, billing model (per-second vs per-hour), minimum commitment, storage pricing, and egress fees. Where available it also shows provider-specific features: CoreWeave's InfiniBand fabric for multi-node training, Lambda Labs' no-egress-fee policy, RunPod's community cloud vs secure cloud tiers, and Vast.ai's marketplace model. Use the comparison to identify the cheapest provider for a specific GPU model, or to evaluate total cost of ownership including storage and networking.
Hyperscalers vs specialist GPU clouds
The GPU cloud market divides into two tiers. Hyperscalers — AWS, Google Cloud, Azure, Oracle Cloud — offer the broadest region coverage, enterprise SLAs, and deep integration with managed services (S3, BigQuery, Azure Blob), but at premium prices: AWS p5.48xlarge (8× H100 SXM) costs $98.32/hr on-demand. Specialist GPU clouds — CoreWeave, Lambda Labs, RunPod, Vast.ai, Nebius AI — offer H100 at $3.50–$5.50/hr with faster provisioning and simpler billing. For pure GPU compute without managed service dependencies, specialist clouds typically offer 60–80% cost savings over hyperscalers.
Spot vs on-demand across providers
Spot (preemptible) GPU availability and pricing varies significantly by provider. CoreWeave offers interruptible instances at 60% below on-demand with 30-second notice. RunPod's spot instances are 40–70% cheaper but can be interrupted with minimal notice. Vast.ai's marketplace model means spot-equivalent pricing is the default, with interruption risk depending on the host machine's utilisation. Lambda Labs does not offer spot instances — all instances are on-demand. AWS and Google Cloud spot/preemptible instances are available but interruption rates for H100 are high during peak demand periods.
Hidden costs: egress, storage, and networking
The hourly GPU rate is only part of the true cost. Egress fees — charged when data leaves the provider's network — are $0.08–$0.09/GB on AWS and Google Cloud, adding up quickly for large model checkpoints and dataset transfers. Lambda Labs and CoreWeave charge no egress fees, which can save thousands of dollars per month for teams moving large datasets. Persistent storage costs $0.02–$0.10/GB/month depending on provider and storage tier. Multi-GPU training jobs using InfiniBand or NVLink fabric may incur additional networking charges on some providers. The comparison tool surfaces these costs where data is available.
Billing models and minimum increments
Billing granularity matters for short-running jobs. Providers billing per-second (RunPod, Vast.ai, Lambda Labs) are significantly cheaper for jobs under one hour — a 20-minute fine-tuning run on RunPod costs one-third of the same run on a provider billing per-hour. CoreWeave bills per-minute with a one-minute minimum. AWS and Google Cloud bill per-second for most instance types with a one-minute minimum. For long-running training jobs (24+ hours), billing granularity is irrelevant — focus on the hourly rate and spot availability instead.
Best providers for H100 training
For large-scale LLM training requiring 8× or more H100 SXM nodes with InfiniBand interconnect, CoreWeave and Lambda Labs are the leading specialist options. CoreWeave offers bare-metal H100 clusters with 3.2 Tbps InfiniBand fabric, per-minute billing, and no egress fees. Lambda Labs offers H100 clusters with 3.2 Tbps InfiniBand at competitive on-demand rates. AWS p5.48xlarge provides 8× H100 SXM with 3.2 Tbps EFA networking but at $98.32/hr — roughly 2× CoreWeave's rate. For single-node H100 jobs, RunPod and Vast.ai offer the lowest on-demand rates.
Best providers for inference serving
Production inference serving requires reliable uptime, autoscaling, and predictable latency — different requirements from training. Modal and Replicate offer serverless GPU inference with automatic scaling to zero, ideal for variable-traffic applications. CoreWeave's Kubernetes-native platform supports custom inference deployments with autoscaling. For managed inference without infrastructure management, Together AI, Groq, and Fireworks AI offer dedicated inference APIs for open-weight models at competitive per-token rates. The right choice depends on whether you need custom model serving or are comfortable with a managed API.
Best providers for budget AI workloads
For cost-sensitive workloads — fine-tuning smaller models, running 7B–13B inference, or experimentation — Vast.ai and RunPod offer the lowest rates. RTX 4090 instances start from $0.44/hr on Vast.ai, making them the most cost-effective option for models that fit in 24GB VRAM. A100 40GB instances are available from $1.50/hr on RunPod. Salad's distributed GPU network offers even lower rates for fault-tolerant batch workloads. For teams with flexible timing, spot instances on any provider can reduce costs by 40–70% versus on-demand.