Track how LLM inference pricing has evolved across providers. Historical token price trends, cost reduction curves, tier breakdowns, and forward-looking analysis — updated regularly.
GPT-4-class models dropped from $30/1M input tokens in mid-2023 to under $5 by mid-2025, driven by efficiency gains and competition from open-weight alternatives.
Llama 3.1 8B and similar models are available for $0.05–$0.10/1M input tokens on Groq, Together AI, and Fireworks — 300× cheaper than frontier models from 2023.
The average supported context window expanded from ~16K tokens in 2023 to 128K–1M tokens by 2025, enabling new long-document and multi-turn use cases.
Groq and Cerebras leveraged custom silicon to push throughput past 1,000 tokens/sec — a 5× improvement over GPU-based inference from 2023.
| Provider / Model | Tier | Input /1M | Output /1M | Cached /1M | TTFT | Throughput | Context | Intel. | Caps | |
|---|---|---|---|---|---|---|---|---|---|---|
Loading pricing data… | ||||||||||
Running your own model on rented GPU compute can be significantly cheaper at scale. A single H100 at ~$2.50/hr can serve ~500K tokens/min of Llama 3.3 70B — compare that against API costs for your expected volume. GPU pricing table and benchmarks page to estimate your self-hosted cost.