GLM 5.2 inference · US-hosted

GLM 5.2 API pricing,
at the full 1M context.

Half the output rate of every provider in the table below, on a dedicated endpoint, with cached input billed at a fraction of list.

50% lower output rate
1M context, served in full
OpenAI-compatible API
Deploys US, VPC, or on-prem
MIT weights, unmodified
Who built the stack
FrontierTrained the largest AI models on the first exascale supercomputer
ACM Gordon Bell finalist, 2024 and 2025
49,152GPUs in record-scale AI training and inference

R&D 100 2025 winner · NVIDIA and national labs, DOE Genesis Mission · Backed by The General Partnership and Ritual Capital.

50%
Lower output rate than every other row in the table below
See the full comparison
GLM 5.2 API pricingCompetitor rows: OpenRouter effective pricing, July 2026
ProviderInput $/1MOutput $/1MCache hit rate
PRIMALABS$0.133$2.2090%
Z.AI$0.346$4.4092.4%
Fireworks$0.421$4.4077.7%
Together AI$0.534$4.4075.9%

Competitor rows are OpenRouter effective pricing, July 2026 — a snapshot, not a contract. PrimaLabs is an independent inference provider and is not affiliated with Z.AI or Zhipu AI; other providers are named for price comparison only. Comparing with DeepSeek V4 Flash? See the DeepSeek table. Page updated 18 August 2026.

Cost estimator

What GLM 5.2 costs on your monthly volume.

$0
Lower per month on PrimaLabs
$0 on PrimaLabs against $0 today · 0% lower

Calculated from the table above at published rates — a first-pass API cost estimate. Dedicated pricing steps down as sustained throughput rises, so this is an upper bound; the benchmark models total cost of ownership and how far the loop can reduce GLM 5.2 API cost and inference spend on your real mix.

The mechanism

Two rates decide a GLM bill: cached input, and output.

Input

Cached input follows the hit rate

A cache hit is a prompt prefix already processed — a system prompt, a codebase, a document set. Z.AI's published GLM-5.2 rate card prices cached input at $0.26 per 1M against $1.40 list, so the effective input rate falls as the hit rate climbs.

Output

Output is never cached

No hit rate discounts a generated token, so the output rate applies to every token, every time. That is why half the output rate — $2.20 against $4.40 — moves a generation-heavy bill more than any input discount.

Warm stack

One dedicated stack, zero routing hops

Your endpoint, your cache, your reservation — AI inference on PrimaLabs dedicated US capacity. Nothing resets the cache by bouncing requests between providers, which is how the hit rate holds at 90%.

Loop

Tuned to the workload, not a benchmark

The learning loop re-tunes batching, scheduling, and cache strategy on live prompts, context lengths, and load. Volume tiers step down as sustained throughput rises.

Coding agents and the 1M window

GLM 5.2 for coding agents: the workload it gets pointed at most.

Z.AI's own developer documentation calls a 90.9% cache hit rate “the average level for coding workloads” — agents replay enormous repeated context, from frontend scaffolding to long-horizon refactors, which is exactly the shape caching rewards and exactly the traffic a dedicated warm stack holds onto.

Cost per task

Agents price out per task, not per token

An AI coding assistant backend or agentic workflow replays the same codebase and system context across many steps, so the margin number is cost per task at your hit rate — with output billed at half rate on every generated token.

Repository scale

The full 1M window, actually served

Repository-scale prompts and long-horizon plans need the whole published window — if you sit under a 256K context API ceiling today, the 1M window changes what fits in one prompt. Worth checking on any quote: the ceiling actually served can sit below the published one.

RAG and documents

Long prefixes, warm cache

Retrieval and document pipelines share long prefixes across requests, so effective cost on a warm dedicated cache diverges sharply from list price — and whole documents fit in context.

Side by side

GLM 5.2 on PrimaLabs vs calling the Z.AI API directly.

Same open model either way — the comparison is the serving layer and the terms. PrimaLabs is an independent GLM 5.2 inference provider and is not affiliated with Z.AI or Zhipu AI.

GLM 5.2 on PrimaLabsZ.AI API, as published
Rate card, per 1M$0.133 input effective at 90% hit rate · $2.20 output$1.40 input list · $0.26 cached · $4.40 output (docs.z.ai, verified August 2026)
ServingDedicated inference on reserved throughput, tuned to your workloadShared API from the model owner's cloud
Where it runsUS cloud, your VPC, on-prem, or customer GPUsZ.AI's infrastructure
Model and contextPublished MIT weights, unmodified, full 1M window servedSame published model and window
Evidence before buyingA workload benchmark on your prompts, in writingPublished rates and docs

Z.AI column: its published rate card and documentation only. Effective-pricing comparison across providers, with hit rates stated, is in the table at the top of this page.

Before you commit

The GLM 5.2 benchmark: throughput, latency, and cost on your prompts, in writing.

Your prompts, your request rate, head to head against the stack you run today — within days, at no charge. Measured at the context lengths you actually send, because long-context latency is where serving layers separate.

Throughput

Tokens per second, and per GPU

Sustained under your concurrency, on a 744B MoE with 40B active per token.

Latency

TTFT and p95, at your context lengths

Time to first token and the tail 95% of requests beat — including at 1M-token prompts, so low latency inference stays predictable under real load.

Cost

Per 1M tokens on your mix

Effective cost on your prompt shapes and cache behaviour, with the serving configuration stated.

Comparison

Comparing GLM 5.2 providers? Compare these eight things.

Shortlisting the best GLM 5.2 API, hosting, or inference provider for 2026 usually starts with quotes from Z.AI (Zhipu AI), Fireworks, Together AI, or an OpenRouter route — and list price is the easiest column to compare and the least predictive of the invoice.

What to compareGLM 5.2 on PrimaLabsAsk your current provider
Output rate per 1M tokens$2.20 — half of every other row above, on every generated tokenWhat is my output rate, and how much of my bill is generation?
Context actually servedThe full published 1M token window, tuned for long-context memory efficiencyWhat ceiling do you actually serve, and does pricing change above a threshold?
Effective input cost$0.133 cache-aware; tiers step down with sustained volumeWhat is my blended input rate after cache, not the list price?
Endpoint type and rate limitsDedicated reserved throughput; the ceiling follows the agreementAm I on a shared pool, and what is the real per-key concurrency ceiling?
Where it runsUS cloud, your VPC, or on-prem in your own data center — NVIDIA or AMDCan this run inside my VPC or on hardware I own, in the region I need?
Data pathKnown, fixed, auditable. No third-party provider touches a requestWhich subprocessors see my prompts, and in which jurisdiction?
Model integrityPublished MIT weights, 744B MoE, 40B active, unmodified; serving config in writingIs the model quantized or modified, and will you state it in writing?
EvidenceA workload benchmark on your prompts, within days, in writingWill you benchmark my traffic, or point me at a leaderboard?

The PrimaLabs column is our commitment. The third column is deliberately a question, not an assertion about another vendor — sourced price rows for named providers are in the table above.

Hosting and deployment

GLM API hosting: dedicated endpoint, VPC, on-prem, or customer GPUs.

Enterprise hosting, deployment, and API pricing questions in one table — dedicated inference on reserved throughput, a managed AI inference service from PrimaLabs, or a deployment service on your own AI infrastructure — one inference solution across all of it.

GLM 5.2 on PrimaLabs
API surfaceOpenAI-compatible: change the base URL and the API key, keep the OpenAI SDK and the rest of the code
Capacity, rate limits, SLAServerless or dedicated. Dedicated is reserved throughput on one warm stack; ceiling and SLA follow the agreement, not a public quota tier
Context servedThe published 1M token window
Model and licenceGLM 5.2 at its published MIT-licensed weights, 744B MoE, 40B active. GLM 5.1 and earlier serve on the same stack
Where it runsUS-based PrimaLabs cloud with sovereign hosting in the USA, your private cloud or VPC, on-premise in your own data center, or a regional deployment — NVIDIA or AMD
Data residencyUS by default. Inside a VPC or on-prem, prompts and completions stay in your compliance boundary
Tenancy and data pathDedicated capacity is your endpoint, your cache, your throughput reservation, on a known, fixed, auditable data path
Reliability record300B+ tokens served across all models, 3.2M requests completed, 0 errors recorded
Continuity and isolation requirementsIf procurement requires single-tenant isolation, multi-region or regional failover, disaster recovery, a private network endpoint, or an air-gapped or hybrid cloud deployment — bring the requirement to the call and it is answered in writing before you commit
Time to productionBenchmark and pricing within days; most teams serve production traffic the same week, with observability and tuning built in
Build versus buy

Already running vLLM or SGLang for GLM, or about to?

On your fleet

Keep the GPUs, drop the tuning work

If you already own or reserve capacity, the same stack deploys onto it — tokens per GPU on hardware you pay for either way, without building the serving layer yourself.

The real cost

1M context is a memory problem

Serving the full window under real concurrency is a continuous memory-efficiency and scheduling problem. A static config decays as prompt shapes drift; the loop absorbs that drift.

Precision

Numeric format is a stated parameter

FP8, INT4, NVFP4 — required or forbidden, it is a deployment decision, stated in writing with the benchmark. Weights stay exactly as published.

Where it fits

Built for workloads where inference is core COGS.

Retention-bound

GLM 5.2 for chatbots and consumer AI

Repeated system prompts and conversation history at high concurrency, on a dedicated reservation rather than a burst quota. Perceived speed drives retention.

Volume-bound

GLM 5.2 for document analysis and batch pipelines

Document analysis, extraction, batch and async inference pipelines where throughput per GPU is the unit economic — and whole documents fit in the served window.

Context-bound

GLM 5.2 for enterprise search and knowledge bases

Enterprise RAG endpoints and knowledge base APIs share long prefixes across requests, so effective cost on a warm dedicated cache diverges sharply from list price.

Same stack, any open model — GLM, Qwen, Kimi, DeepSeek, and MiniMax serve on one PrimaLabs control plane, so the LLM API behind a workload is a deployment choice.

How it starts

Traffic profile in. Benchmark, pricing, and a live endpoint out.

Step 01

Share the traffic profile

Prompt shapes, context length distribution, request rate, and the cache behaviour you see today.

Step 02

Head-to-head benchmark

GLM 5.2 against your current stack — throughput, latency, and cost per million tokens, at your context lengths, in writing.

Step 03

Dedicated pricing

PrimaLabs dedicated pricing: a cache-aware cost model, no routing markup, volume tiers that step down as sustained throughput rises.

Step 04

Change the base URL

OpenAI-compatible, so migration is a drop-in swap: base URL, API key, done — whether you are coming off a closed-model API or an earlier GLM release. Most teams serve production traffic the same week.

FAQ

GLM API pricing, hosting, and deployment questions

What does the GLM 5.2 API cost, and how is pricing structured?
Per million tokens on dedicated capacity: $0.133 input and $2.20 output in the table above — half the output rate of every other row. Pricing is cache-aware and volume tiers step down as sustained throughput rises.
What is a cache hit rate, and why does it decide the input bill?
A cache hit is prompt work the server has already done and can reuse. Z.AI's published GLM-5.2 rate card prices cached input at $0.26 against $1.40 list, so the hit rate sets the effective input rate. Output tokens are never cached.
How do we get GLM API access, and what changes in our code?
You get your own endpoint on an OpenAI-compatible API. Change the base URL and the API key, keep the rest of the code — OpenAI SDK integration is the whole migration.
Can we run GLM 5.2 in our own VPC, or on our own GPUs?
Yes. Endpoints deploy in the US-based PrimaLabs cloud, a private cloud or customer VPC, or on-prem on customer GPUs, NVIDIA or AMD — one sovereign stack wherever it runs.
Is this cheaper than self-hosting GLM 5.2?
That depends on utilisation and engineering time, and it is what the benchmark answers. If you already own GPUs, the same stack deploys onto your fleet, so the question becomes tokens per GPU rather than rent versus own.
Do the optimizations change model output?
No. GLM 5.2 runs at its published MIT-licensed weights. Optimization lives in the serving layer — engine, kernels, precision, scheduling — and the exact serving configuration goes in writing alongside the benchmark.
What does the benchmark cost, and what do we get back?
Nothing. Share the traffic profile and you get a head-to-head workload benchmark against your current stack plus dedicated GLM 5.2 pricing with a cache-aware cost model, on real prompts, within days, in writing. There is no self-serve tier.
Is GLM 5.2 served at the full 1M token context?
Yes — the published 1M token window, with the serving layer tuned for long-context memory efficiency under real traffic. Worth checking on any quote: the ceiling actually served can sit below the published window.
Is GLM 5.2 a good backend for a coding agent?
It is the workload GLM 5.2 gets pointed at most. Z.AI's own developer documentation describes a 90.9% cache hit rate as the average level for coding workloads — repetitive codebase context is exactly the shape caching rewards, and agents price out per task, not per token.
Do you quantize GLM 5.2, or change serving precision?
Weights are untouched. If your evaluation requires or forbids a numeric format — FP8, INT4, NVFP4 — that is a deployment parameter, stated in writing with the benchmark.
Which GLM generations do you serve?
Any open model on the same stack — GLM, Qwen, Kimi, DeepSeek, MiniMax. If production is on GLM 5.1, GLM 4.6 or 4.5, or an older ChatGLM release, bring the version to the call and the serving and migration path is confirmed before you commit.
Is PrimaLabs affiliated with Z.AI or Zhipu AI?
No. PrimaLabs is an independent GLM 5.2 inference provider. Z.AI (Zhipu AI) is the model's first-party API provider; its rates appear above for price comparison only.
Why does the output rate matter more than the input rate on GLM 5.2?
Output tokens are never cached, so no hit rate discounts them — the output rate applies to every generated token. That is why half the output rate ($2.20 against $4.40) moves a generation-heavy bill more than any input discount.
What about high availability, multi-region, or air-gapped requirements?
Bring the requirement to the call — continuity, isolation, and deployment requirements are scoped and answered in writing before you commit. Reliability record across all models to date: 300B+ tokens, 3.2M requests, 0 errors.

Get GLM 5.2 pricing for your workload

300B+ tokens served, 3.2M requests completed, 0 errors recorded — across all models, on US-based dedicated capacity. Share the traffic profile and get dedicated GLM 5.2 pricing plus a workload benchmark on real prompts, within days, in writing.

Prefer email? [email protected] · Or request API access and we reach out to you.

Request API access

Dedicated capacity is scoped to your workload, so access is provisioned after a short engineering review — not self-served.

Used only to reply to you. Privacy Policy