Compute Comparison
vs
All providers →

Voyage AI vs Z.AI: Token Pricing, Speed & Intelligence

Full comparison of Voyage AI and Z.AI — live token pricing, latency, throughput, context window, strengths, weaknesses, and best use cases. Updated July 2026.

Voyage AI

State-of-the-art embedding and reranking models

Voyage AI specialises in embedding and reranking models for retrieval-augmented generation (RAG) and semantic search. Voyage 3.5 and its variants consistently top the MTEB leaderboard for retrieval quality.

RAG pipelinesSemantic searchDocument retrievalReranking
Proprietary models

Z.AI

GLM frontier models with 1M context

Z.AI (formerly Zhipu AI) develops the GLM series of large language models. GLM-5.2 supports a 1M token context window and is designed for enterprise-grade chat, coding, and long-document tasks.

ChatCodingLong-document analysisEnterprise AI
Proprietary models

Key metrics

Cheapest input ($/1M)

Cheapest output ($/1M)

Peak throughput

Best latency (TTFT)

Intelligence score

Context window

Live token pricing

Strengths & weaknesses

Voyage AI

Top MTEB leaderboard performance
Multimodal embedding support
Very competitive pricing
Embeddings and reranking only — no chat models
Smaller ecosystem than OpenAI

Z.AI

1M token context window
Strong Chinese and English bilingual performance
Enterprise-grade reliability
Smaller international developer community
Fewer third-party integrations than OpenAI

Key differentiators

Voyage AI

Voyage 3.5 Lite offers top-tier retrieval quality at just $0.02/1M tokens — the most cost-effective high-quality embedding available.

Z.AI

GLM-5.2 offers a 1M token context window at $1.11/1M input tokens, making it one of the most cost-effective long-context models available.

Frequently asked questions

Voyage AI FAQs

What is Voyage AI used for?

Voyage AI provides embedding and reranking models for RAG pipelines, semantic search, and document retrieval. It does not offer chat or text generation models.

How does Voyage AI compare to OpenAI embeddings?

Voyage 3.5 consistently outperforms OpenAI text-embedding-3-large on MTEB benchmarks while being significantly cheaper. It is the preferred choice for production RAG systems.

Z.AI FAQs

What is GLM-5.2?

GLM-5.2 is the latest model in Zhipu AI's GLM series, supporting a 1M token context window. It is designed for long-document analysis, coding, and enterprise chat applications.

How does Z.AI compare to other Chinese LLM providers?

Z.AI's GLM models compete with Alibaba's Qwen and Baidu's ERNIE series. GLM-5.2 stands out for its 1M context window and competitive pricing.

Provider resources

Voyage AIState-of-the-art embedding and reranking models

Voyage AI specialises in embedding and reranking models for retrieval-augmented generation (RAG) and semantic search. Voyage 3.5 and its variants consistently top the MTEB leaderboard for retrieval quality.

Voyage 3.5 Lite offers top-tier retrieval quality at just $0.02/1M tokens — the most cost-effective high-quality embedding available.

Z.AIGLM frontier models with 1M context

Z.AI (formerly Zhipu AI) develops the GLM series of large language models. GLM-5.2 supports a 1M token context window and is designed for enterprise-grade chat, coding, and long-document tasks.

GLM-5.2 offers a 1M token context window at $1.11/1M input tokens, making it one of the most cost-effective long-context models available.

Key strengths compared

Voyage AI

  • Top MTEB leaderboard performance
  • Multimodal embedding support
  • Very competitive pricing

Z.AI

  • 1M token context window
  • Strong Chinese and English bilingual performance
  • Enterprise-grade reliability

Provider category context

Voyage AI is a inference api, founded in 2023. Z.AI is a frontier lab, founded in 2019. Voyage AI as an inference API provider hosts open-weight models — typically offering lower prices for equivalent capability tiers. Z.AI as a frontier lab trains and serves proprietary models with capabilities not available elsewhere.

How to choose between them

Choose Voyage AI if you need top mteb leaderboard performance. Choose Z.AI if you need 1m token context window. For high-volume production workloads, run a cost comparison using the token pricing table above with your actual prompt/completion token ratio — the cheapest provider depends heavily on your input-to-output token ratio.