DeepSeek has cut the price of cached input tokens on its V4-Pro model by roughly 90%, dropping the effective cost of repeated system prompts, tool definitions and retrieval context to near-zero and intensifying a price war that increasingly hinges on agentic workloads.
May 24, 2026 · DeepSeek V4-Pro
The 90% Cache Cut That Redraws the AI Price War
DeepSeek slashed cached-input pricing tenfold, dropping the cost of repeated system prompts, tool definitions and retrieval context to near-zero — and making cache economics, not sticker rates, the decisive factor for agents at scale.
90%
cut to cache-hit price (tenfold reduction)
75%
cut to headline V4-Pro API rates
Cost at scale · 1B tokens · 90% cache-hit rate
V4-Pro runs an agent for a fraction of Grok 4.5
A ~12× spread — cache pricing, not headline rates, now determines the true cost of running an agent.
V4-Pro price per million tokens
Blended ~$0.18/M tokens (7:2:1 weighting) — repeated 100K+ prefixes become economically trivial to resend.
Why it matters
In agentic and RAG pipelines the same system prompt, tool definitions and context circulate through dozens or hundreds of calls. That boilerplate now costs almost nothing — cost pressure shifts to fresh, uncached tokens.
The caveats
Surge pricing during peak Beijing hours can double rates. Savings depend on cache-hit rates — poorly structured prompts that break the cached prefix erode the advantage.
The open question
Will xAI and others answer with cache cuts of their own?
Grok 4.5 lists a ~75% cache discount but still trails on effective cost. If rivals hold, the price of reusing context may become the decisive factor in which model teams pick to power their agents.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…