GLM 5.2 API pricing,
at the full 1M context.
Half the output rate of every provider in the table below, on a dedicated endpoint, with cached input billed at a fraction of list.
R&D 100 2025 winner · NVIDIA and national labs, DOE Genesis Mission · Backed by The General Partnership and Ritual Capital.
| Provider | Input $/1M | Output $/1M | Cache hit rate |
|---|---|---|---|
| PRIMALABS | $0.133 | $2.20 | 90% |
| Z.AI | $0.346 | $4.40 | 92.4% |
| Fireworks | $0.421 | $4.40 | 77.7% |
| Together AI | $0.534 | $4.40 | 75.9% |
Competitor rows are OpenRouter effective pricing, July 2026 — a snapshot, not a contract. PrimaLabs is an independent inference provider and is not affiliated with Z.AI or Zhipu AI; other providers are named for price comparison only. Comparing with DeepSeek V4 Flash? See the DeepSeek table. Page updated 18 August 2026.
What GLM 5.2 costs on your monthly volume.
Calculated from the table above at published rates — a first-pass API cost estimate. Dedicated pricing steps down as sustained throughput rises, so this is an upper bound; the benchmark models total cost of ownership and how far the loop can reduce GLM 5.2 API cost and inference spend on your real mix.
Two rates decide a GLM bill: cached input, and output.
Cached input follows the hit rate
A cache hit is a prompt prefix already processed — a system prompt, a codebase, a document set. Z.AI's published GLM-5.2 rate card prices cached input at $0.26 per 1M against $1.40 list, so the effective input rate falls as the hit rate climbs.
Output is never cached
No hit rate discounts a generated token, so the output rate applies to every token, every time. That is why half the output rate — $2.20 against $4.40 — moves a generation-heavy bill more than any input discount.
One dedicated stack, zero routing hops
Your endpoint, your cache, your reservation — AI inference on PrimaLabs dedicated US capacity. Nothing resets the cache by bouncing requests between providers, which is how the hit rate holds at 90%.
Tuned to the workload, not a benchmark
The learning loop re-tunes batching, scheduling, and cache strategy on live prompts, context lengths, and load. Volume tiers step down as sustained throughput rises.
| Cache hit rate | 0% | 50% | 75% | 90% | 100% |
|---|---|---|---|---|---|
| Effective input $/1M | $1.40 | $0.83 | $0.55 | $0.37 | $0.26 |
GLM 5.2 for coding agents: the workload it gets pointed at most.
Z.AI's own developer documentation calls a 90.9% cache hit rate “the average level for coding workloads” — agents replay enormous repeated context, from frontend scaffolding to long-horizon refactors, which is exactly the shape caching rewards and exactly the traffic a dedicated warm stack holds onto.
Agents price out per task, not per token
An AI coding assistant backend or agentic workflow replays the same codebase and system context across many steps, so the margin number is cost per task at your hit rate — with output billed at half rate on every generated token.
The full 1M window, actually served
Repository-scale prompts and long-horizon plans need the whole published window — if you sit under a 256K context API ceiling today, the 1M window changes what fits in one prompt. Worth checking on any quote: the ceiling actually served can sit below the published one.
Long prefixes, warm cache
Retrieval and document pipelines share long prefixes across requests, so effective cost on a warm dedicated cache diverges sharply from list price — and whole documents fit in context.
GLM 5.2 on PrimaLabs vs calling the Z.AI API directly.
Same open model either way — the comparison is the serving layer and the terms. PrimaLabs is an independent GLM 5.2 inference provider and is not affiliated with Z.AI or Zhipu AI.
| GLM 5.2 on PrimaLabs | Z.AI API, as published | |
|---|---|---|
| Rate card, per 1M | $0.133 input effective at 90% hit rate · $2.20 output | $1.40 input list · $0.26 cached · $4.40 output (docs.z.ai, verified August 2026) |
| Serving | Dedicated inference on reserved throughput, tuned to your workload | Shared API from the model owner's cloud |
| Where it runs | US cloud, your VPC, on-prem, or customer GPUs | Z.AI's infrastructure |
| Model and context | Published MIT weights, unmodified, full 1M window served | Same published model and window |
| Evidence before buying | A workload benchmark on your prompts, in writing | Published rates and docs |
Z.AI column: its published rate card and documentation only. Effective-pricing comparison across providers, with hit rates stated, is in the table at the top of this page.
The GLM 5.2 benchmark: throughput, latency, and cost on your prompts, in writing.
Your prompts, your request rate, head to head against the stack you run today — within days, at no charge. Measured at the context lengths you actually send, because long-context latency is where serving layers separate.
Tokens per second, and per GPU
Sustained under your concurrency, on a 744B MoE with 40B active per token.
TTFT and p95, at your context lengths
Time to first token and the tail 95% of requests beat — including at 1M-token prompts, so low latency inference stays predictable under real load.
Per 1M tokens on your mix
Effective cost on your prompt shapes and cache behaviour, with the serving configuration stated.
Comparing GLM 5.2 providers? Compare these eight things.
Shortlisting the best GLM 5.2 API, hosting, or inference provider for 2026 usually starts with quotes from Z.AI (Zhipu AI), Fireworks, Together AI, or an OpenRouter route — and list price is the easiest column to compare and the least predictive of the invoice.
| What to compare | GLM 5.2 on PrimaLabs | Ask your current provider |
|---|---|---|
| Output rate per 1M tokens | $2.20 — half of every other row above, on every generated token | What is my output rate, and how much of my bill is generation? |
| Context actually served | The full published 1M token window, tuned for long-context memory efficiency | What ceiling do you actually serve, and does pricing change above a threshold? |
| Effective input cost | $0.133 cache-aware; tiers step down with sustained volume | What is my blended input rate after cache, not the list price? |
| Endpoint type and rate limits | Dedicated reserved throughput; the ceiling follows the agreement | Am I on a shared pool, and what is the real per-key concurrency ceiling? |
| Where it runs | US cloud, your VPC, or on-prem in your own data center — NVIDIA or AMD | Can this run inside my VPC or on hardware I own, in the region I need? |
| Data path | Known, fixed, auditable. No third-party provider touches a request | Which subprocessors see my prompts, and in which jurisdiction? |
| Model integrity | Published MIT weights, 744B MoE, 40B active, unmodified; serving config in writing | Is the model quantized or modified, and will you state it in writing? |
| Evidence | A workload benchmark on your prompts, within days, in writing | Will you benchmark my traffic, or point me at a leaderboard? |
The PrimaLabs column is our commitment. The third column is deliberately a question, not an assertion about another vendor — sourced price rows for named providers are in the table above.
GLM API hosting: dedicated endpoint, VPC, on-prem, or customer GPUs.
Enterprise hosting, deployment, and API pricing questions in one table — dedicated inference on reserved throughput, a managed AI inference service from PrimaLabs, or a deployment service on your own AI infrastructure — one inference solution across all of it.
| GLM 5.2 on PrimaLabs | |
|---|---|
| API surface | OpenAI-compatible: change the base URL and the API key, keep the OpenAI SDK and the rest of the code |
| Capacity, rate limits, SLA | Serverless or dedicated. Dedicated is reserved throughput on one warm stack; ceiling and SLA follow the agreement, not a public quota tier |
| Context served | The published 1M token window |
| Model and licence | GLM 5.2 at its published MIT-licensed weights, 744B MoE, 40B active. GLM 5.1 and earlier serve on the same stack |
| Where it runs | US-based PrimaLabs cloud with sovereign hosting in the USA, your private cloud or VPC, on-premise in your own data center, or a regional deployment — NVIDIA or AMD |
| Data residency | US by default. Inside a VPC or on-prem, prompts and completions stay in your compliance boundary |
| Tenancy and data path | Dedicated capacity is your endpoint, your cache, your throughput reservation, on a known, fixed, auditable data path |
| Reliability record | 300B+ tokens served across all models, 3.2M requests completed, 0 errors recorded |
| Continuity and isolation requirements | If procurement requires single-tenant isolation, multi-region or regional failover, disaster recovery, a private network endpoint, or an air-gapped or hybrid cloud deployment — bring the requirement to the call and it is answered in writing before you commit |
| Time to production | Benchmark and pricing within days; most teams serve production traffic the same week, with observability and tuning built in |
Already running vLLM or SGLang for GLM, or about to?
Keep the GPUs, drop the tuning work
If you already own or reserve capacity, the same stack deploys onto it — tokens per GPU on hardware you pay for either way, without building the serving layer yourself.
1M context is a memory problem
Serving the full window under real concurrency is a continuous memory-efficiency and scheduling problem. A static config decays as prompt shapes drift; the loop absorbs that drift.
Numeric format is a stated parameter
FP8, INT4, NVFP4 — required or forbidden, it is a deployment decision, stated in writing with the benchmark. Weights stay exactly as published.
Built for workloads where inference is core COGS.
GLM 5.2 for chatbots and consumer AI
Repeated system prompts and conversation history at high concurrency, on a dedicated reservation rather than a burst quota. Perceived speed drives retention.
GLM 5.2 for document analysis and batch pipelines
Document analysis, extraction, batch and async inference pipelines where throughput per GPU is the unit economic — and whole documents fit in the served window.
GLM 5.2 for enterprise search and knowledge bases
Enterprise RAG endpoints and knowledge base APIs share long prefixes across requests, so effective cost on a warm dedicated cache diverges sharply from list price.
Same stack, any open model — GLM, Qwen, Kimi, DeepSeek, and MiniMax serve on one PrimaLabs control plane, so the LLM API behind a workload is a deployment choice.
Traffic profile in. Benchmark, pricing, and a live endpoint out.
Share the traffic profile
Prompt shapes, context length distribution, request rate, and the cache behaviour you see today.
Head-to-head benchmark
GLM 5.2 against your current stack — throughput, latency, and cost per million tokens, at your context lengths, in writing.
Dedicated pricing
PrimaLabs dedicated pricing: a cache-aware cost model, no routing markup, volume tiers that step down as sustained throughput rises.
Change the base URL
OpenAI-compatible, so migration is a drop-in swap: base URL, API key, done — whether you are coming off a closed-model API or an earlier GLM release. Most teams serve production traffic the same week.
GLM API pricing, hosting, and deployment questions
What does the GLM 5.2 API cost, and how is pricing structured?
What is a cache hit rate, and why does it decide the input bill?
How do we get GLM API access, and what changes in our code?
Can we run GLM 5.2 in our own VPC, or on our own GPUs?
Is this cheaper than self-hosting GLM 5.2?
Do the optimizations change model output?
What does the benchmark cost, and what do we get back?
Is GLM 5.2 served at the full 1M token context?
Is GLM 5.2 a good backend for a coding agent?
Do you quantize GLM 5.2, or change serving precision?
Which GLM generations do you serve?
Is PrimaLabs affiliated with Z.AI or Zhipu AI?
Why does the output rate matter more than the input rate on GLM 5.2?
What about high availability, multi-region, or air-gapped requirements?
Get GLM 5.2 pricing for your workload
300B+ tokens served, 3.2M requests completed, 0 errors recorded — across all models, on US-based dedicated capacity. Share the traffic profile and get dedicated GLM 5.2 pricing plus a workload benchmark on real prompts, within days, in writing.
Prefer email? [email protected] · Or request API access and we reach out to you.