Compute Comparison

GPU Price History Charts & Trends — H100, A100 & RTX Daily Rates Across 94+ Providers

Loading price history…

How H100 prices have changed since 2024

NVIDIA H100 80GB SXM on-demand prices have fallen roughly 40–50% from early 2024 to mid-2026. In January 2024, H100 on-demand rates ranged from $4.50/hr (specialist providers) to $12.00/hr (AWS p5 on-demand). By July 2026, the range had compressed to $2.49/hr (Lambda Labs, Fluidstack) to $5.89/hr (AWS). The primary driver is supply expansion: CoreWeave, Lambda Labs, and a wave of new specialist providers brought large H100 clusters online through 2024–2025, breaking the supply bottleneck that kept prices elevated after the H100's 2022 launch. Spot prices have fallen even further — from $3.00+/hr in early 2024 to $1.60–$1.89/hr on Vast.ai and RunPod in 2026.

A100 price trajectory and displacement

A100 80GB on-demand prices have declined from $2.50–$4.00/hr in early 2024 to $1.29–$3.67/hr by mid-2026, as H100 supply growth has pushed providers to discount older A100 inventory. The A100 is increasingly displaced in new deployments — its 312 TFLOPS FP16 and 2 TB/s HBM2e bandwidth are roughly 3× lower than the H100's 989 TFLOPS and 3.35 TB/s. For inference workloads where the A100's specs are sufficient, the price decline makes it an attractive option: at $1.29/hr on RunPod, the A100 80GB offers 80GB of HBM2e at a lower cost than any H100 configuration. A100 40GB instances are available from $0.89/hr, making them the most cost-effective option for 7B–13B model inference where 40GB of VRAM is sufficient.

RTX 4090 pricing and the consumer GPU market

RTX 4090 cloud pricing has remained relatively stable at $0.34–$0.79/hr since mid-2024, making it the most cost-effective option for 7B–13B model inference and fine-tuning. Unlike data center GPUs, RTX 4090 supply is constrained by consumer GPU production cycles rather than hyperscaler procurement, so prices have not fallen as sharply as H100 or A100. The 24GB GDDR6X VRAM fits most 7B models at FP16 and 13B models at INT8, and the 1,008 GB/s bandwidth is competitive with older A100 40GB configurations. RunPod community cloud and Vast.ai marketplace consistently offer the lowest RTX 4090 rates — often $0.34–$0.44/hr for spot instances.

What drives GPU cloud price changes

GPU cloud prices are shaped by four forces. Supply expansion is the dominant driver: when CoreWeave, Lambda Labs, or a new entrant brings a large H100 cluster online, on-demand rates across the market typically drop 5–15% within weeks as providers compete for workloads. Demand spikes — driven by major model releases, research paper drops, or enterprise AI adoption waves — temporarily push spot prices up and reduce availability. New GPU generations (H200, Blackwell B200) create a cascading effect: as newer hardware becomes available, providers discount the previous generation to maintain utilization. Finally, hyperscaler pricing (AWS, Google Cloud, Azure) sets a ceiling — specialist providers consistently undercut hyperscaler rates by 40–60% for equivalent hardware, which anchors the upper end of the market.

GPU price forecast: H2 2026 and 2027

H100 on-demand prices are projected to continue declining through H2 2026 and into 2027. As NVIDIA Blackwell (B200) supply increases and H100 transitions to second-generation status, specialist providers are expected to offer H100 80GB SXM at $1.99–$2.49/hr by end of 2026 and $1.49–$1.99/hr by mid-2027. Spot rates may fall below $1.00/hr on marketplace platforms as H100 inventory grows relative to demand. A100 80GB is projected to reach $0.99–$1.29/hr on-demand by end of 2026 as it is further displaced by H100. These projections assume continued supply expansion and no major demand shock — a large-scale AI training run or enterprise adoption wave could temporarily reverse the trend. Use the price history charts above to track actual rates against these projections.

How to use GPU price history for procurement decisions

The 90-day price history charts are most useful for three procurement decisions. First, timing on-demand purchases: if a GPU's price has been declining steadily over 30 days, waiting 2–4 weeks before committing to a long-running job may save 5–15%. Second, evaluating reserved instance commitments: if a provider's on-demand rate has been stable for 60+ days, a 3-month or 6-month reserved commitment at a 20–30% discount is likely to pay off. Third, identifying spot price windows: spot prices often dip during weekends and off-peak hours — the history charts reveal these patterns, letting you schedule batch training jobs during low-cost windows. Filter by GPU model and provider to isolate the specific price series relevant to your workload.

Comparing hyperscalers vs specialist providers over time

The price history data reveals a persistent and widening gap between hyperscaler (AWS, Google Cloud, Azure) and specialist provider (Lambda Labs, CoreWeave, RunPod, Fluidstack) GPU rates. In early 2024, AWS p5 H100 on-demand was roughly 2.5× more expensive than Lambda Labs. By mid-2026, that gap has grown to 2.4× ($5.89/hr vs $2.49/hr). Hyperscalers compete on ecosystem integration, compliance certifications, and SLA guarantees rather than raw GPU price — their rates have declined more slowly because enterprise customers pay a premium for managed networking, IAM, and support. For pure compute workloads without hyperscaler ecosystem dependencies, specialist providers consistently offer 40–60% lower rates for equivalent hardware. The history charts make this gap visible over time, helping teams quantify the cost of staying on a hyperscaler versus migrating to a specialist provider.