DeepSeek V4 Flash API pricing,
decided by cache hit rate.
Cached input bills at a fraction of list price — your effective rate follows your hit rate, not a sticker price.
R&D 100 2025 winner · NVIDIA and national labs, DOE Genesis Mission · Backed by The General Partnership and Ritual Capital.
| Provider | Input $/1M | Output $/1M | Cache hit rate |
|---|---|---|---|
| PRIMALABS | $0.016 | $0.266 | 90% |
| DeepSeek | $0.031 | $0.279 | 79.6% |
| NovitaAI | $0.053 | $0.279 | 78.0% |
| Baidu Qianfan | $0.037 | $0.181 | 73.6% |
| Fireworks | $0.100 | $0.279 | 35.2% |
Competitor rows are OpenRouter effective pricing, July 2026 — a snapshot, not a contract. Deployment options — US, VPC, on-prem — are detailed below. PrimaLabs is an independent inference provider; other providers are named for price comparison only. Comparing with GLM 5.2? See the GLM 5.2 table. Page updated 18 August 2026.
What DeepSeek V4 Flash costs on your monthly volume.
Calculated from the table above at published rates — a first-pass API cost estimate. Dedicated pricing steps down as sustained throughput rises, so this is an upper bound; total cost of ownership is settled by the benchmark.
Your effective rate is your hit rate, priced.
A cache hit is reused work
A hit is a prompt prefix the server has already processed — a system prompt, a codebase, a document set — billed at a fraction of list. DeepSeek's own rate card prices the spread at 50×: $0.0028 per 1M cached against $0.14 uncached.
One dedicated stack, zero routing hops
Your endpoint, your cache, your reservation — AI inference on PrimaLabs dedicated US capacity. Nothing resets the cache by bouncing requests between providers, which is how the hit rate stays at 90% instead of decaying.
Tuned to the workload, not a benchmark
The learning loop watches live prompts, context lengths, and load, then re-tunes batching, scheduling, and cache strategy for that exact traffic — custom inference in the literal sense. Serving gets cheaper the longer it runs.
Tiers step down as volume rises
Cache-aware pricing plus volume tiers that fall with sustained throughput, so the cache hit cost savings land directly on the bill. No routing markup on top.
| Cache hit rate | 0% | 50% | 75% | 90% | 100% |
|---|---|---|---|---|---|
| Effective input $/1M | $0.140 | $0.071 | $0.037 | $0.017 | $0.0028 |
The DeepSeek benchmark: throughput, latency, and cost on your prompts, in writing.
High-performance inference is measured, not asserted: your prompts, your request rate, head to head against the stack you run today — within days, at no charge. There is no trial and no card; this is the pre-sales artifact a demo-led motion runs on. A trial on a cold shared endpoint tells you what a cold endpoint costs; this tells you what your workload costs.
Tokens per second, and per GPU
Sustained under your concurrency — the unit economic of a high-volume pipeline.
TTFT and p95 under load
Time to first token and the tail 95% of requests beat — measured under load, not on an idle box, because low latency inference is proven under concurrency or not at all.
Per 1M tokens on your mix
Effective cost on your actual prompt shapes and cache behaviour, with the serving configuration stated.
Comparing DeepSeek API providers? Compare these eight things.
Whether you are shortlisting AI inference service providers broadly or holding a quote from DeepSeek's official API, NovitaAI, Baidu Qianfan, Fireworks, or an OpenRouter route — list price is the easiest column to compare and the least predictive of the invoice.
| What to compare | DeepSeek V4 Flash on PrimaLabs | Ask your current provider |
|---|---|---|
| Effective cost per 1M tokens | $0.016 input, $0.266 output, cache-aware, tiers step down with volume | What is my blended rate after cache, not the list price? |
| Cache hit rate on your traffic | 90% on a warm dedicated stack, actively managed upward | What hit rate do I actually get, and is it measured for my tenant? |
| Endpoint type and rate limits | Dedicated reserved throughput; the ceiling follows the agreement | Am I on a shared pool, and what is the real quota ceiling? |
| Routing | Zero hops. Every request hits the same warm stack | How many providers can serve my request, and does config vary between them? |
| Where it runs | US cloud, your private cloud or VPC, or on-prem on your GPUs, NVIDIA or AMD | Can this run inside my VPC or on hardware I own? |
| Data path | Known, fixed, auditable. No third-party provider touches a request | Which subprocessors see my prompts? |
| Model integrity | Published weights, 284B MoE, unmodified; serving config in writing on request | Is the model quantized or modified, and will you state it in writing? |
| Evidence | A workload benchmark on your prompts, within days, in writing | Will you benchmark my traffic, or point me at a leaderboard? |
The PrimaLabs column is our commitment. The third column is deliberately a question, not an assertion about another vendor — sourced price rows for named providers are in the table above.
DeepSeek API hosting: dedicated endpoint, US data residency, VPC or on-prem.
Enterprise hosting, deployment, and API pricing questions in one table — dedicated inference on reserved throughput, managed inference services from PrimaLabs, or a deployment service on your own AI infrastructure — one inference solution across all of it.
| DeepSeek V4 Flash on PrimaLabs | |
|---|---|
| API surface | OpenAI-compatible: change the base URL and the API key, keep the OpenAI SDK and the rest of the code |
| Capacity, rate limits, SLA | Serverless or dedicated. Dedicated is reserved throughput on one warm stack; ceiling and SLA follow the agreement, not a public quota tier |
| Where it runs | US-based PrimaLabs cloud with sovereign hosting in the USA, your private cloud or VPC, on-premise in your own data center, or a regional deployment — on PrimaLabs capacity or customer GPUs: inference on custom hardware, NVIDIA or AMD |
| Data residency | US by default. Inside a VPC or on-prem, prompts and completions stay in your compliance boundary |
| Tenancy and isolation | Dedicated capacity is your endpoint, your cache, your throughput reservation. If procurement requires isolated or single-tenant hosting, bring the requirement to the call — it is answered in writing before you commit |
| Data path | Known, fixed, auditable — no routing layer, no third-party provider |
| Model integrity | Weights exactly as published; optimization lives in the serving layer, with the configuration supplied in writing when your evaluation needs it |
| Reliability record | 300B+ tokens served across all models, 3.2M requests completed, 0 errors recorded |
| Time to production | Benchmark and pricing within days; most teams serve production traffic the same week, with observability and tuning built in |
Dedicated endpoint vs shared serverless vs self-hosted DeepSeek.
Serverless inference on PrimaLabs covers spiky, low-volume work; dedicated model inference on reserved throughput is the production default. The table is the honest version of that choice.
| Dedicated endpoint | Shared serverless | Self-hosted | |
|---|---|---|---|
| Serving layer run by | PrimaLabs, tuned to your traffic | The provider, tuned for the whole pool | Your team |
| Rate-limit ceiling | Your reserved throughput, per agreement | A public quota tier | Whatever your fleet sustains |
| Cache behaviour | Your cache, warm, managed toward 90% | Shared — evicted by other tenants' traffic | Yours to engineer and keep tuned |
| Cost shape | Cache-aware per 1M, tiers step down with volume | List price per 1M | GPU capacity plus engineering time |
| Fits best | Production traffic with an SLA and a bill worth optimizing | Prototypes and spiky, low-volume work | Teams with an idle owned fleet — see below |
Keep the GPUs, drop the tuning work
If you already own or reserve GPU capacity, the same stack deploys onto it. The question stops being rent versus own and becomes tokens per GPU on hardware you pay for either way.
A 284B MoE is a continuous tuning problem
Mixture of experts routes each token through a small subset of the network, so packing and scheduling decide throughput. Standing the model up is a weekend; holding a 90% hit rate as prompt shapes drift is a permanent engineering commitment.
Bake it off on your own cluster
Benchmark head to head against your current config. If your stack wins, you keep a documented baseline and lost nothing.
Built for workloads where inference is core COGS.
DeepSeek for coding agents and voice AI
Both stack time to first token on every step. High hit rates on repeated codebase or session context cut cost per session, and zero routing hops keep low latency inference predictable at the tail.
DeepSeek for RAG and document processing
Retrieval and document processing pipelines share long prefixes across requests — the shape caching rewards most — served at the published 1M token context. Inference for batch workloads prices the same way: throughput per GPU.
DeepSeek for high-volume chatbots and support AI
Repeated system prompts and history at high concurrency, on a dedicated reservation rather than a burst quota. Perceived speed drives retention.
Same stack, any open model — DeepSeek, GLM, Qwen, Kimi, and MiniMax serve on one PrimaLabs control plane, so the LLM API behind a workload is a deployment choice.
Traffic profile in. Benchmark, pricing, and a live endpoint out.
Share the traffic profile
Prompt shapes, context length distribution, request rate, and the cache behaviour you see today.
Head-to-head benchmark
Throughput, latency, and cost per million tokens against your current stack — on real prompts, within days, in writing.
Dedicated pricing
PrimaLabs dedicated pricing: a cache-aware cost model, no routing markup, volume tiers that step down as sustained throughput rises.
Change the base URL
OpenAI-compatible, so migration is a drop-in swap: base URL, API key, done. Most teams serve production traffic the same week.
DeepSeek API pricing, access, and deployment questions
How is DeepSeek V4 Flash API pricing structured?
What is a cache hit rate, and why does it decide the bill?
How do we get DeepSeek API access, and what changes in our code?
Can we run DeepSeek V4 Flash in our own VPC, or on our own GPUs?
Is this cheaper than self-hosting DeepSeek V4 Flash?
Do the optimizations change model output?
What does the benchmark cost, and what do we get back?
Where does the $0.016 effective input figure come from?
Which DeepSeek generation do you serve, and can we migrate from R1, Reasoner, or V3.1?
Does DeepSeek V4 Flash run at long context for RAG and document work?
Are there rate limits on a dedicated endpoint?
Is the data path really free of third parties?
How fast is time to production?
What performance numbers do you report, and on whose traffic?
Get DeepSeek V4 Flash pricing for your workload
300B+ tokens served, 3.2M requests completed, 0 errors recorded — across all models, on US-based dedicated capacity. Share the traffic profile and get dedicated pricing plus a workload benchmark on real prompts, within days, in writing.
Prefer email? [email protected] · Or request API access and we reach out to you.