Cheaper than any provider
on OpenRouter.
Open models on dedicated, sovereign serving. One stack instead of a router, with a learning loop that keeps cache hit rates high, lowering effective token costs while keeping performance at peak.
| Provider | Input $/1M | Output $/1M | Cache hit rate |
|---|---|---|---|
| PRIMALABS | $0.016 | $0.266 | 90% |
| DeepSeek | $0.031 | $0.279 | 79.6% |
| NovitaAI | $0.053 | $0.279 | 78.0% |
| Fireworks | $0.100 | $0.279 | 35.2% |
| Provider | Input $/1M | Output $/1M | Cache hit rate |
|---|---|---|---|
| PRIMALABS | $0.133 | $2.20 | 90% |
| Z.AI | $0.346 | $4.40 | 92.4% |
| Fireworks | $0.421 | $4.40 | 77.7% |
| Together AI | $0.534 | $4.40 | 75.9% |
Static stacks tune once. Ours learns every workload.
Tuned to the workload, not the benchmark
The learning loop observes live prompts, context lengths, and load, then re-tunes batching, scheduling, and cache strategy for that exact traffic. Every model on the stack, its serving gets faster and cheaper the longer it runs.
Hit rate is where the bill is won
Cached input tokens bill at a fraction of list price. One warm, dedicated stack keeps hit rates high where routed serverless traffic fragments them. Effective cost per token drops below any sticker price.
Dedicated capacity, fully controlled
Reserved throughput on one sovereign stack, in the PrimaLabs cloud, a customer VPC, or on customer GPUs, NVIDIA or AMD. Known data path, predictable latency, cost that falls with scale.
PrimaLabs versus OpenRouter
| PrimaLabs | OpenRouter | |
|---|---|---|
| Infrastructure model | Sovereign, dedicated capacity | Serverless router over shared third-party capacity |
| Cache behavior | One warm stack, hit rate managed by the learning loop | Cache state fragments across whichever provider gets the request |
| Workload optimization | Learns each workload and re-tunes continuously | None. Inherits whatever the routed provider tuned once |
| Data path | Known, fixed, auditable. In-VPC and on-prem available | Traverses 60+ external providers. No sovereignty |
| Capacity | Reserved, guaranteed throughput | Serverless, best effort |
Common questions
What does sovereign mean here?
Why highlight cache hit rate?
Can existing OpenRouter code migrate?
What does customization mean in practice?
See the numbers on real traffic
Tell us about the workload and the models it runs. Get dedicated pricing and a cache-aware cost model within days.