DeepSeek V4 Flash inference · US-hosted

DeepSeek V4 Flash API pricing,
decided by cache hit rate.

Cached input bills at a fraction of list price — your effective rate follows your hit rate, not a sticker price.

48% lower input rate
Dedicated endpoint, zero hops
OpenAI-compatible API
US data residency
Deploys VPC or on-prem
Benchmark in writing
Who built the stack
FrontierTrained the largest AI models on the first exascale supercomputer
ACM Gordon Bell finalist, 2024 and 2025
49,152GPUs in record-scale AI training and inference

R&D 100 2025 winner · NVIDIA and national labs, DOE Genesis Mission · Backed by The General Partnership and Ritual Capital.

48%
Lower effective input rate than DeepSeek's own API, at the rates below
See the full comparison
DeepSeek V4 Flash API pricingCompetitor rows: OpenRouter effective pricing, July 2026
ProviderInput $/1MOutput $/1MCache hit rate
PRIMALABS$0.016$0.26690%
DeepSeek$0.031$0.27979.6%
NovitaAI$0.053$0.27978.0%
Baidu Qianfan$0.037$0.18173.6%
Fireworks$0.100$0.27935.2%

Competitor rows are OpenRouter effective pricing, July 2026 — a snapshot, not a contract. Deployment options — US, VPC, on-prem — are detailed below. PrimaLabs is an independent inference provider; other providers are named for price comparison only. Comparing with GLM 5.2? See the GLM 5.2 table. Page updated 18 August 2026.

Cost estimator

What DeepSeek V4 Flash costs on your monthly volume.

$0
Lower per month on PrimaLabs
$0 on PrimaLabs against $0 today · 0% lower

Calculated from the table above at published rates — a first-pass API cost estimate. Dedicated pricing steps down as sustained throughput rises, so this is an upper bound; total cost of ownership is settled by the benchmark.

The mechanism

Your effective rate is your hit rate, priced.

Cache

A cache hit is reused work

A hit is a prompt prefix the server has already processed — a system prompt, a codebase, a document set — billed at a fraction of list. DeepSeek's own rate card prices the spread at 50×: $0.0028 per 1M cached against $0.14 uncached.

Warm stack

One dedicated stack, zero routing hops

Your endpoint, your cache, your reservation — AI inference on PrimaLabs dedicated US capacity. Nothing resets the cache by bouncing requests between providers, which is how the hit rate stays at 90% instead of decaying.

Loop

Tuned to the workload, not a benchmark

The learning loop watches live prompts, context lengths, and load, then re-tunes batching, scheduling, and cache strategy for that exact traffic — custom inference in the literal sense. Serving gets cheaper the longer it runs.

Tiers

Tiers step down as volume rises

Cache-aware pricing plus volume tiers that fall with sustained throughput, so the cache hit cost savings land directly on the bill. No routing markup on top.

Before you commit

The DeepSeek benchmark: throughput, latency, and cost on your prompts, in writing.

High-performance inference is measured, not asserted: your prompts, your request rate, head to head against the stack you run today — within days, at no charge. There is no trial and no card; this is the pre-sales artifact a demo-led motion runs on. A trial on a cold shared endpoint tells you what a cold endpoint costs; this tells you what your workload costs.

Throughput

Tokens per second, and per GPU

Sustained under your concurrency — the unit economic of a high-volume pipeline.

Latency

TTFT and p95 under load

Time to first token and the tail 95% of requests beat — measured under load, not on an idle box, because low latency inference is proven under concurrency or not at all.

Cost

Per 1M tokens on your mix

Effective cost on your actual prompt shapes and cache behaviour, with the serving configuration stated.

Comparison

Comparing DeepSeek API providers? Compare these eight things.

Whether you are shortlisting AI inference service providers broadly or holding a quote from DeepSeek's official API, NovitaAI, Baidu Qianfan, Fireworks, or an OpenRouter route — list price is the easiest column to compare and the least predictive of the invoice.

What to compareDeepSeek V4 Flash on PrimaLabsAsk your current provider
Effective cost per 1M tokens$0.016 input, $0.266 output, cache-aware, tiers step down with volumeWhat is my blended rate after cache, not the list price?
Cache hit rate on your traffic90% on a warm dedicated stack, actively managed upwardWhat hit rate do I actually get, and is it measured for my tenant?
Endpoint type and rate limitsDedicated reserved throughput; the ceiling follows the agreementAm I on a shared pool, and what is the real quota ceiling?
RoutingZero hops. Every request hits the same warm stackHow many providers can serve my request, and does config vary between them?
Where it runsUS cloud, your private cloud or VPC, or on-prem on your GPUs, NVIDIA or AMDCan this run inside my VPC or on hardware I own?
Data pathKnown, fixed, auditable. No third-party provider touches a requestWhich subprocessors see my prompts?
Model integrityPublished weights, 284B MoE, unmodified; serving config in writing on requestIs the model quantized or modified, and will you state it in writing?
EvidenceA workload benchmark on your prompts, within days, in writingWill you benchmark my traffic, or point me at a leaderboard?

The PrimaLabs column is our commitment. The third column is deliberately a question, not an assertion about another vendor — sourced price rows for named providers are in the table above.

Hosting and deployment

DeepSeek API hosting: dedicated endpoint, US data residency, VPC or on-prem.

Enterprise hosting, deployment, and API pricing questions in one table — dedicated inference on reserved throughput, managed inference services from PrimaLabs, or a deployment service on your own AI infrastructure — one inference solution across all of it.

DeepSeek V4 Flash on PrimaLabs
API surfaceOpenAI-compatible: change the base URL and the API key, keep the OpenAI SDK and the rest of the code
Capacity, rate limits, SLAServerless or dedicated. Dedicated is reserved throughput on one warm stack; ceiling and SLA follow the agreement, not a public quota tier
Where it runsUS-based PrimaLabs cloud with sovereign hosting in the USA, your private cloud or VPC, on-premise in your own data center, or a regional deployment — on PrimaLabs capacity or customer GPUs: inference on custom hardware, NVIDIA or AMD
Data residencyUS by default. Inside a VPC or on-prem, prompts and completions stay in your compliance boundary
Tenancy and isolationDedicated capacity is your endpoint, your cache, your throughput reservation. If procurement requires isolated or single-tenant hosting, bring the requirement to the call — it is answered in writing before you commit
Data pathKnown, fixed, auditable — no routing layer, no third-party provider
Model integrityWeights exactly as published; optimization lives in the serving layer, with the configuration supplied in writing when your evaluation needs it
Reliability record300B+ tokens served across all models, 3.2M requests completed, 0 errors recorded
Time to productionBenchmark and pricing within days; most teams serve production traffic the same week, with observability and tuning built in
Build versus buy

Dedicated endpoint vs shared serverless vs self-hosted DeepSeek.

Serverless inference on PrimaLabs covers spiky, low-volume work; dedicated model inference on reserved throughput is the production default. The table is the honest version of that choice.

Dedicated endpointShared serverlessSelf-hosted
Serving layer run byPrimaLabs, tuned to your trafficThe provider, tuned for the whole poolYour team
Rate-limit ceilingYour reserved throughput, per agreementA public quota tierWhatever your fleet sustains
Cache behaviourYour cache, warm, managed toward 90%Shared — evicted by other tenants' trafficYours to engineer and keep tuned
Cost shapeCache-aware per 1M, tiers step down with volumeList price per 1MGPU capacity plus engineering time
Fits bestProduction traffic with an SLA and a bill worth optimizingPrototypes and spiky, low-volume workTeams with an idle owned fleet — see below
On your fleet

Keep the GPUs, drop the tuning work

If you already own or reserve GPU capacity, the same stack deploys onto it. The question stops being rent versus own and becomes tokens per GPU on hardware you pay for either way.

The real cost

A 284B MoE is a continuous tuning problem

Mixture of experts routes each token through a small subset of the network, so packing and scheduling decide throughput. Standing the model up is a weekend; holding a 90% hit rate as prompt shapes drift is a permanent engineering commitment.

Proof first

Bake it off on your own cluster

Benchmark head to head against your current config. If your stack wins, you keep a documented baseline and lost nothing.

Where it fits

Built for workloads where inference is core COGS.

Latency-bound

DeepSeek for coding agents and voice AI

Both stack time to first token on every step. High hit rates on repeated codebase or session context cut cost per session, and zero routing hops keep low latency inference predictable at the tail.

Context-bound

DeepSeek for RAG and document processing

Retrieval and document processing pipelines share long prefixes across requests — the shape caching rewards most — served at the published 1M token context. Inference for batch workloads prices the same way: throughput per GPU.

Volume-bound

DeepSeek for high-volume chatbots and support AI

Repeated system prompts and history at high concurrency, on a dedicated reservation rather than a burst quota. Perceived speed drives retention.

Same stack, any open model — DeepSeek, GLM, Qwen, Kimi, and MiniMax serve on one PrimaLabs control plane, so the LLM API behind a workload is a deployment choice.

How it starts

Traffic profile in. Benchmark, pricing, and a live endpoint out.

Step 01

Share the traffic profile

Prompt shapes, context length distribution, request rate, and the cache behaviour you see today.

Step 02

Head-to-head benchmark

Throughput, latency, and cost per million tokens against your current stack — on real prompts, within days, in writing.

Step 03

Dedicated pricing

PrimaLabs dedicated pricing: a cache-aware cost model, no routing markup, volume tiers that step down as sustained throughput rises.

Step 04

Change the base URL

OpenAI-compatible, so migration is a drop-in swap: base URL, API key, done. Most teams serve production traffic the same week.

FAQ

DeepSeek API pricing, access, and deployment questions

How is DeepSeek V4 Flash API pricing structured?
Per million tokens on dedicated capacity: $0.016 input and $0.266 output in the table above. Pricing is cache-aware, volume tiers step down as sustained throughput rises, and there is no routing markup.
What is a cache hit rate, and why does it decide the bill?
A cache hit is prompt work the server has already done and can reuse. Cached input bills at a fraction of list price — DeepSeek's own rate card prices a hit at $0.0028 against $0.14 for a miss — so the hit rate, not the list rate, sets the invoice.
How do we get DeepSeek API access, and what changes in our code?
You get your own endpoint on an OpenAI-compatible API. Change the base URL and the API key, keep the rest of the code and the OpenAI SDK you already use — full API compatibility, a drop-in replacement.
Can we run DeepSeek V4 Flash in our own VPC, or on our own GPUs?
Yes. Endpoints deploy in the US-based PrimaLabs cloud, a private cloud or customer VPC, or on-prem on customer GPUs, NVIDIA or AMD — one sovereign stack wherever it runs.
Is this cheaper than self-hosting DeepSeek V4 Flash?
That depends on utilisation and engineering time, and it is what the benchmark answers. If you already own GPUs, the same stack deploys onto your fleet, so the question becomes tokens per GPU rather than rent versus own.
Do the optimizations change model output?
No. Weights stay exactly as published. All gains come from the host-level serving layer — batching, scheduling, cache, and memory management — and the exact serving configuration goes in writing alongside the benchmark.
What does the benchmark cost, and what do we get back?
Nothing. Share the traffic profile and you get a head-to-head workload benchmark against your current stack plus dedicated pricing with a cache-aware cost model, on real prompts, within days, in writing. There is no self-serve tier; this is the pre-sales artifact.
Where does the $0.016 effective input figure come from?
Effective input cost is hit rate × cached rate + (1 − hit rate) × list rate. The table states each provider's observed hit rate alongside its effective price, with competitor rows from OpenRouter effective pricing, July 2026.
Which DeepSeek generation do you serve, and can we migrate from R1, Reasoner, or V3.1?
The same stack serves any open model across DeepSeek, GLM, Qwen, Kimi, and MiniMax. If production is on R1, Reasoner, or V3.1, bring the version to the call and the migration path to V4 Flash is confirmed before you commit.
Does DeepSeek V4 Flash run at long context for RAG and document work?
Yes — served at the published 1M token context, with the serving layer tuned for long-context memory efficiency. Long shared prefixes are where cache economics matter most.
Are there rate limits on a dedicated endpoint?
Dedicated capacity means reserved throughput rather than best-effort serverless from a shared pool, so the ceiling follows the throughput in your agreement, not a public quota tier.
Is the data path really free of third parties?
Yes. The data path is known, fixed, and auditable. No request touches a third-party provider, and there is no routing layer between your application and the model.
How fast is time to production?
Benchmark and pricing within days of a traffic profile. Because migration is a base URL and key swap, most teams are serving production traffic the same week.
What performance numbers do you report, and on whose traffic?
Tokens per second and throughput per GPU, time to first token, p95 latency under your concurrency, and cost per million tokens — measured on your prompts against your current stack, returned in writing. Never a public leaderboard position.

Get DeepSeek V4 Flash pricing for your workload

300B+ tokens served, 3.2M requests completed, 0 errors recorded — across all models, on US-based dedicated capacity. Share the traffic profile and get dedicated pricing plus a workload benchmark on real prompts, within days, in writing.

Prefer email? [email protected] · Or request API access and we reach out to you.

Request API access

Dedicated capacity is scoped to your workload, so access is provisioned after a short engineering review — not self-served.

Used only to reply to you. Privacy Policy