Groq vs Together AI: Token Pricing, Speed & Intelligence
Full comparison of Groq and Together AI — live token pricing, latency, throughput, context window, strengths, weaknesses, and best use cases. Updated July 2026.
Groq
LPU-powered inference — the fastest tokens per second available
Groq runs custom Language Processing Units (LPUs) that deliver dramatically higher throughput than GPU-based inference — Llama 3.3 70B reaches 750+ tokens/second on Groq, versus 100–200 on typical GPU providers. Ideal for latency-sensitive applications, real-time chat, and high-volume batch workloads.
Together AI
Open-source model hosting with competitive inference pricing
Together AI specialises in hosting open-weight models including the full Llama family, Mixtral, and DeepSeek variants. They offer live pricing via their public API and support fine-tuning workflows. A popular choice for teams that want open-source flexibility without managing their own GPU infrastructure.
Key metrics
—
—
—
—
—
—
—
—
—
—
—
—
Live token pricing
Strengths & weaknesses
Groq
Together AI
Key differentiators
Groq's custom LPU chips deliver 750+ tokens/sec on Llama 3.3 70B — 4–5× faster than any GPU-based provider.
The broadest open-weight model catalog with fine-tuning support — ideal for teams that need model customisation without self-hosting.
Frequently asked questions
Groq FAQs
How fast is Groq inference?
Groq delivers 750+ tokens/second on Llama 3.3 70B and 1,200+ tokens/second on Llama 3.1 8B. This is 4–5× faster than typical GPU-based providers, making it ideal for real-time applications.
How much does Groq cost?
Llama 3.3 70B costs $0.59/1M input and $0.79/1M output tokens. Llama 3.1 8B is just $0.05/$0.08 per 1M tokens — among the cheapest options for a capable open-weight model.
What is a Groq LPU?
A Language Processing Unit (LPU) is Groq's custom silicon designed specifically for sequential token generation. Unlike GPUs which are optimised for parallel matrix operations, LPUs excel at the autoregressive decoding step that dominates LLM inference latency.
Together AI FAQs
What models does Together AI support?
Together AI hosts 100+ open-weight models including the full Llama 3.x family (8B, 70B, 405B), Mixtral, DeepSeek R1, Qwen, and many others. They also support custom fine-tuned model deployment.
How much does Together AI cost?
Llama 3.3 70B costs $0.88/1M tokens (input and output). Llama 3.1 405B is $3.50/1M tokens. Smaller models like Llama 3.2 11B Vision start at $0.18/1M tokens.
Does Together AI support fine-tuning?
Yes. Together AI offers supervised fine-tuning for Llama and other open-weight models. You can upload training data, run fine-tuning jobs, and deploy the resulting model via their inference API.