Compute Comparison

LLM Provider Inference Comparison — Speed, Intelligence & Token Cost Across 18+ Providers

GPU Model Comparison, Price History Charts & LLM Inference Pricing — Compute Comparison

GPU Model Comparison

Select up to 4 GPUs to compare specs and live pricing side-by-side

Select GPUs to Compare (3/4)

H100 80GB

GH100 · TSMC 4N

Hopper

VRAM

80 GB HBM3

Bandwidth

3350 GB/s

TDP

700 W

Released

2022

Interconnect

NVLink 4.0 / PCIe 5.0

Flagship data center GPU. Transformer Engine with FP8 support. SXM5 form factor for max bandwidth.

A100 80GB

GA100 · TSMC 7nm

Ampere

VRAM

80 GB HBM2e

Bandwidth

2000 GB/s

TDP

400 W

Released

2020

Interconnect

NVLink 3.0 / PCIe 4.0

Workhorse of the AI era. Widely available, mature software support, excellent price/performance.

RTX 4090

AD102 · TSMC 4N

Ada Lovelace

VRAM

24 GB GDDR6X

Bandwidth

1008 GB/s

TDP

450 W

Released

2022

Interconnect

PCIe 4.0

Consumer GPU with surprisingly strong FP32. No NVLink. Best $/TFLOP for budget workloads.

Performance Comparison

FP16 PerformanceTFLOPS
H100 80GB
2.0K
A100 80GB
312
RTX 4090
165.2
BF16 PerformanceTFLOPS
H100 80GB
2.0K
A100 80GB
312
RTX 4090
165.2
INT8 PerformanceTOPS
H100 80GB
4.0K
A100 80GB
624
RTX 4090
330.3
VRAMGB
H100 80GB
80
A100 80GB
80
RTX 4090
24
Memory BandwidthGB/s
H100 80GB
3.4K
A100 80GB
2.0K
RTX 4090
1.0K
FP32 PerformanceTFLOPS
H100 80GB
67
A100 80GB
19.5
RTX 4090
82.6
Power Draw (TDP)W
H100 80GB
700
A100 80GB
400
RTX 4090
450

Full Spec Sheet

SpecH100 80GBA100 80GBRTX 4090
ArchitectureGH100GA100AD102
GenerationHopperAmpereAda Lovelace
Process NodeTSMC 4NTSMC 7nmTSMC 4N
Transistors80B54.2B76.3B
VRAM80 GB80 GB24 GB
VRAM TypeHBM3HBM2eGDDR6X
Mem Bandwidth3350 GB/s2000 GB/s1008 GB/s
FP3267 TFLOPS19.5 TFLOPS82.6 TFLOPS
FP161979 TFLOPS312 TFLOPS165.2 TFLOPS
BF161979 TFLOPS312 TFLOPS165.2 TFLOPS
FP83958 TFLOPS
INT83958 TOPS624 TOPS330.3 TOPS
NVLink BW900 GB/s600 GB/s
InterconnectNVLink 4.0 / PCIe 5.0NVLink 3.0 / PCIe 4.0PCIe 4.0
TDP700 W400 W450 W
Released202220202022

Best Use Cases

H100 80GB

  • LLM training
  • Large-scale inference
  • Scientific HPC

A100 80GB

  • ML training
  • Large model inference
  • HPC workloads

RTX 4090

  • Fine-tuning small models
  • Inference up to 13B
  • Cost-sensitive workloads

How providers are ranked

Each LLM inference provider is scored across four dimensions: intelligence (derived from MMLU, HumanEval, and MATH benchmark scores), throughput (tokens/sec at p50 load), latency (time-to-first-token in milliseconds), and cost efficiency (intelligence score per dollar per 1M tokens). The radar chart visualises all four axes simultaneously, making it easy to identify providers that excel in one area — Groq for raw speed, OpenAI for intelligence, Together AI and Fireworks AI for cost efficiency on open-weight models.

Frontier vs open-weight providers

The LLM inference market divides into two tiers. Frontier model providers — OpenAI, Anthropic, Google — offer proprietary models (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro) with the highest intelligence scores but at premium prices ($2.50–$15/1M input tokens). Open-weight inference providers — Groq, Together AI, Fireworks AI, Cerebras, Nebius AI — serve Llama 3.1, Mistral, Qwen, and DeepSeek models at 10–100× lower cost, often with higher throughput due to specialised hardware like Groq's LPU or Cerebras's wafer-scale chips.

Choosing the right provider for your workload

Provider selection depends on your workload's dominant constraint. Latency-sensitive applications (chatbots, copilots, real-time assistants) should prioritise TTFT — Groq delivers sub-100ms TTFT for Llama 3.1 70B. Throughput-heavy pipelines (batch document processing, RAG indexing) benefit from high tokens/sec — Cerebras and Fireworks AI excel here. Cost-optimised workloads should compare intelligence-per-dollar: DeepSeek V3 at $0.27/1M input tokens delivers GPT-4-class reasoning at a fraction of OpenAI's price. Use the LLM ROI Calculator to model your specific token volume and input/output ratio.

Understanding throughput and latency trade-offs

High throughput and low latency are often in tension. Groq's LPU architecture achieves 500–800 tokens/sec for Llama 3.1 70B with sub-100ms TTFT, but only supports a subset of open-weight models. OpenAI and Anthropic prioritise reliability and model quality — their TTFT is typically 300–800ms for frontier models under normal load, with throughput of 60–120 tokens/sec. For production systems, measure p50 and p99 latency under your expected concurrent request load, not just single-request benchmarks. Provider performance degrades significantly at high concurrency.

Token pricing trends in 2026

LLM inference prices have fallen 90%+ since 2023 across all tiers. GPT-4o is now $2.50/1M input tokens — down from $30/1M for the original GPT-4. Claude 3.5 Haiku costs $0.80/1M input tokens. Open-weight models on inference APIs are cheaper still: Llama 3.1 8B from $0.06/1M on Groq, Mistral 7B from $0.07/1M on Together AI. The efficient tier (Llama 3.1 70B, Mistral Large, Qwen 2.5 72B) now delivers near-frontier intelligence at $0.20–$0.90/1M input tokens, making it the default choice for most production workloads where GPT-4o-level quality is not strictly required.