Compute Comparison
Inference APIOpen Weights

Groq

LPU-powered inference — the fastest tokens per second available

Groq runs custom Language Processing Units (LPUs) that deliver dramatically higher throughput than GPU-based inference — Llama 3.3 70B reaches 750+ tokens/second on Groq, versus 100–200 on typical GPU providers. Ideal for latency-sensitive applications, real-time chat, and high-volume batch workloads.

Groq's custom LPU chips deliver 750+ tokens/sec on Llama 3.3 70B — 4–5× faster than any GPU-based provider.

SpeedReal-time chatVoice AICost-efficiencyBatch processing
Mountain View, CA
Founded 2016
groq.comOfficial pricing pageDocumentation
Strengths
  • 750+ tokens/sec on Llama 3.3 70B — fastest GPU-class inference
  • Sub-100ms time-to-first-token for real-time applications
  • Very competitive pricing on open-weight models
  • OpenAI-compatible API
  • Free tier available
Limitations
  • Limited model selection vs. Together AI or Fireworks
  • No vision model support on most models
  • No fine-tuning capability

All Models

Sort by:
Loading models…

Groq — Frequently Asked Questions

Community Reviews

Loading reviews…

All LLM Providers

Full comparison table

Groq LLM pricing overview

Groq is an inference API provider that hosts open-weight and third-party large language models. Groq runs custom Language Processing Units (LPUs) that deliver dramatically higher throughput than GPU-based inference — Llama 3.3 70B reaches 750+ tokens/second on Groq, versus 100–200 on typical GPU providers. Ideal for latency-sensitive applications, real-time chat, and high-volume batch workloads. Common use cases include Speed, Real-time chat, Voice AI, Cost-efficiency. Headquartered in Mountain View, CA, founded 2016. All models are accessible via a REST API compatible with standard OpenAI-style request formats, enabling drop-in integration with most LLM frameworks and orchestration tools.

Groq vs other LLM providers

Groq competes with OpenAI, Anthropic, Google, Groq, Together AI, Mistral, Cohere, and other inference API providers across dimensions of price, throughput, latency, context window, and model intelligence. Groq's custom LPU chips deliver 750+ tokens/sec on Llama 3.3 70B — 4–5× faster than any GPU-based provider. Use the LLM provider comparison tool to see Groq token pricing, latency, and throughput side-by-side with any other provider. The full LLM pricing table shows all providers ranked by input token cost, output token cost, and throughput in a single sortable view.

Why choose Groq?

Groq's key strengths are: 750+ tokens/sec on Llama 3.3 70B — fastest GPU-class inference; Sub-100ms time-to-first-token for real-time applications; Very competitive pricing on open-weight models. Limitations to consider: Limited model selection vs. Together AI or Fireworks; No vision model support on most models. For teams running high-volume inference workloads, prompt caching and batch API endpoints can reduce effective input token costs by 50–90% — check the context window cost guide for a full breakdown of caching economics.

Understanding Groq token pricing

Groq charges separately for input (prompt) and output (completion) tokens, priced per 1M tokens in USD. Output tokens are typically 3–5× more expensive than input tokens due to the compute cost of autoregressive generation. Prices shown are sourced from Groq's public pricing page and updated daily. Need help estimating your spend? Read the LLM API cost calculator guide — it covers tokens, context windows, prompt caching, and batch discounts with worked examples for RAG, chat history, and document processing workloads.

Groq context window and model capabilities

Context window size directly affects both capability and cost — every token in the context window is charged as an input token. For RAG and document processing workloads, longer context windows enable richer retrieval but increase per-call costs proportionally. Prompt caching — where supported — stores the KV state of repeated prefixes and charges 75–90% less for cache hits, making it the most impactful cost optimization for applications with consistent system prompts or retrieved documents. See the LLM context window cost guide for a full analysis of how context length affects your API bill.

Self-hosted vs managed inference: when Groq makes sense

Managed inference APIs like Groq eliminate infrastructure overhead — no GPU provisioning, driver management, or model serving stack to maintain. The trade-off is cost at scale: a single H100 at ~$2.50/hr can serve ~500K tokens/min of Llama 3.3 70B, which at Groq API rates would cost significantly more per token. The break-even point depends on your request volume, latency requirements, and engineering capacity. For teams processing fewer than ~10M tokens/day, managed APIs are almost always cheaper when total cost of ownership is considered. Above that threshold, self-hosted inference on rented GPU compute typically wins on unit economics. Read the cheapest GPU cloud guide for a full break-even analysis. Historical Groq token price data is available in the LLM price history charts.