DeepSeek V4 pricing, verified: the 4x discrepancy trackers miss
TL;DR: DeepSeek V4 Pro costs $0.435 per 1M cache-miss input tokens and $0.87 per 1M output tokens. V4 Flash costs $0.14 input and $0.28 output. Trackers including BenchLM still quote $1.74/$3.48, a clean 4x too high, because DeepSeek deleted the discount expiry date on 22 May 2026 and made the promo price permanent, confirmed by Reuters and InfoWorld. Peak-hour pricing at 2x is documented on the official pricing page but not yet in effect, despite GA coverage claiming otherwise. Verified against live docs 2026-08-05 at 20:14 CEST.
What you’re looking at is the result of research I finished at 20:14 Copenhagen time on 5 August 2026, after checking every load-bearing number against both official and independent sources. I found one 4x discrepancy between DeepSeek’s live pricing and the figure some trackers still carry.
What’s Inside
What DeepSeek V4 Pro and V4 Flash cost per 1M tokens as of 5 August 2026, checked twice in one day
Why some trackers quote $1.74 and $3.48 when your bill will say $0.435 and $0.87
What happened on 22 May 2026 that nobody’s changelog explains properly
Whether peak-hour pricing is live yet, and why GA coverage says yes while the docs say no
Which peak hours look like in Copenhagen, New York, and San Francisco time
What broke on 24 July 2026 if your code still says
deepseek-chat
The verified pricing table, August 2026
Per 1M tokens, USD. I checked every number against the DeepSeek pricing docs on 2026-08-05.
Two traps are waiting inside this very tidy table.
The 1M context is shared between input and output. It is not additive. We don’t get 1M input tokens plus another 384K output tokens tucked into the glove compartment. We get 1M total, with up to 384K available for output.
The Pro cache-hit price is not a typo. DeepSeek priced cached Pro input at under 1% of cache-miss input. If a workflow hits the cache consistently, Pro gets very cheap.
Compare DeepSeek with what you pay per 1M tokens across your real traffic mix. Pro can handle the difficult work at Flash-adjacent economics when the cache-hit rate does its job.
My PM read: V4 Flash at $0.14 input and $0.28 output, with a 1M context window and 2500-request concurrency, is the most aggressively priced production inference tier I have verified in 2026.
Subscribe to Product with Attitude for verified AI pricing snapshots, discrepancy logs, and builder-side critical AI literacy every week.
The V4 Pro 4x Discrepancy, Explained
DeepSeek’s V4 Pro promotion never expired. On 22 May 2026, DeepSeek deleted the discount’s expiry date from its pricing page, making $0.435 input and $0.87 output the permanent list price instead of a temporary 75% cut. I spent an hour convinced this was undocumented. It is documented extensively, and I was searching the wrong words.
Here is the receipt, from a company statement quoted by InfoWorld on 25 May 2026:
The Deepseek V4 Pro model API pricing will be officially adjusted to 1/4 of the original price after the 75% discount promotion ends on 2026/05/31 15:59 UTC.
Reuters ran it as a straight news item: DeepSeek “will make permanent a 75% price cut on its flagship V4-Pro artificial intelligence model, keeping prices at a quarter of their original level, the company said in a statement on Saturday.” Chinese-language coverage adds that DeepSeek announced the permanence on its official X account (MoneyDJ via TTV). Greyhound Research’s Sanchit Vir Gogia gave InfoWorld the analyst reading: the cut is “permanent rather than promotional,” an efficiency gain passed through rather than a sale.
So the 4x drift has a mundane cause. Trackers froze different moments.
Verdent still lists $1.74 and $3.48, sourced from the official page “as of April 28, 2026,” which was accurate on that date. BenchLM also carries the higher numbers despite a 31 July timestamp. CloudZero hedges by showing both, labelling $1.74/$3.48 the “standard rate” and $0.435/$0.87 a “75% promo,” which was true for nine days and has been wrong since 22 May. BenchLM, DevTk, MorphLLM, Coworker AI, and Codersera match the live page.
Your bill follows the live pricing page on the day of the API call. Today that means the lower number. Worth knowing if you’re forecasting a budget.
The dated timeline nobody keeps in one place:
Peak and Off-Peak Pricing Is Documented But Not Live
DeepSeek plans to charge 2x its regular rate during Beijing business hours, and the effective date has not been announced. The official page, re-read at 20:14 CEST today, says the API “will soon adopt” the policy and that the effective date “will be subject to the official announcement.” Quoting it directly:
“During peak hours, prices will be 2x the regular prices, applicable to all billing items. Peak hours: 9:00–12:00 and 14:00–18:00 (Beijing Time, UTC+8) daily.”
“All billing items” means all of them. Cache-hit input, cache-miss input, and output.
Here is discrepancy number two, and it runs the opposite direction from the first one. Several GA write-ups state that peak pricing shipped with the mid-July general availability release, and two of them add that weekends stay off-peak (MACGPU, Zenn). The live page says neither. It says pending, and it says daily. I get mildly furious about this category of error, because a scheduling decision made on the strength of “weekends are off-peak” costs real money if the docs are the thing that turns out to be right.
Treat peak pricing as announced, unpriced, and daily until DeepSeek says otherwise.
For those of us who do not schedule our lives in Beijing time
Tip: Set background workers against Beijing time. A small cron decision can keep every scheduled run outside the surge window.
The overnight batch prompt:
Audit every scheduled DeepSeek API job in this repo. Convert each cron schedule to Beijing time (UTC+8) and flag any run that falls inside 09:00–12:00 or 14:00–18:00. Output a table of job name, current schedule, Beijing-time window, and a suggested off-peak replacement schedule.What Broke on 24 July 2026
If your code still calls deepseek-chat or deepseek-reasoner, it stopped working. DeepSeek permanently retired both legacy model aliases on 24 July 2026 at 15:59 UTC, and calls using them now fail. I found this buried three sections deep in a GA changelog, which feels like the wrong place for a hard cutoff.
The replacements:
deepseek-chatbecomesdeepseek-v4-flashin non-thinking mode, a drop-in swap for most chat trafficdeepseek-reasonerbecomesdeepseek-v4-flashin thinking mode, ordeepseek-v4-prowhen the task justifies it
Grep your codebase for both strings (Zenn GA breakdown, MACGPU).
The Pricing Sub-Page That Now Returns 404
DeepSeek’s /quick_start/pricing-details-usd sub-page now returns “Page Not Found,” and it used to list retired pre-V4 models that contradicted the main table. One source of contradictory model info is gone, which I count as a win.
Every old bookmark, citation, and AI training snapshot pointing at that URL now dead-ends, which I count as the bill for the win. The current models are deepseek-v4-flash and deepseek-v4-pro, and the only pricing URL worth bookmarking is the primary pricing page.
FAQ
How much does DeepSeek V4 Flash cost per 1M tokens?
I checked the live page because copied pricing tables age with the dignity of supermarket sushi. DeepSeek V4 Flash costs $0.0028 per 1M cached input tokens, $0.14 per 1M cache-miss input tokens, and $0.28 per 1M output tokens.
Corroborated by CloudZero, Verdent, and MorphLLM.
How much does DeepSeek V4 Pro cost per 1M tokens?
DeepSeek V4 Pro costs $0.003625 cache-hit input, $0.435 cache-miss input, and $0.87 output per 1M tokens, and this is the number I recheck before every budget forecast. The undiscounted launch rates of $1.74 and $3.48 survive on Verdent and in CloudZero‘s “standard rate” column. They are historical. Billing follows the live price on the day of the call.
Is the DeepSeek V4 Pro price cut permanent?
Yes, and I got this wrong in my first draft of this page. DeepSeek made the 75% V4 Pro discount permanent on 22 May 2026 by deleting the expiry date rather than letting the promotion lapse on 31 May, confirmed by Reuters and InfoWorld. The “75% promo” framing that several trackers still carry describes a nine-day window that closed in May.
What context window and max output does DeepSeek V4 support?
DeepSeek V4 Flash and V4 Pro each support a 1M-token context window with a maximum output of 384K tokens, sharing one 1M budget. The numbers look more generous than they are. In plain English, 1M input plus 384K output is not available. The combined total cannot exceed 1M (DeepSeek pricing docs, Reapi).
Is DeepSeek peak-hour pricing in effect yet?
DeepSeek's peak-hour pricing at 2x is documented but not yet in effect, with the effective date still unannounced as of 5 August 2026. Peak hours are listed as 09:00–12:00 and 14:00–18:00 Beijing time, daily. Some GA coverage claims it went live in mid-July and that weekends are exempt (MACGPU). The official page says neither.
Which DeepSeek model names still work?
Grep first, read later: only deepseek-v4-flash and deepseek-v4-pro are live, because deepseek-chat and deepseek-reasoner retired permanently on 24 July 2026 at 15:59 UTC. Use deepseek-v4-flash in non-thinking mode as the drop-in replacement for deepseek-chat (Zenn).
What are DeepSeek V4’s concurrency limits?
This is the line I hand straight to whoever owns capacity planning: DeepSeek allows 2500 concurrent requests for V4 Flash and 500 for V4 Pro. The limits apply at account level regardless of how many API keys we create, and exceeding them returns HTTP 429 (DeepSeek rate limit docs).
Does DeepSeek support the Anthropic API format?
DeepSeek provides an Anthropic-compatible endpoint at api.deepseek.com/anthropic and an OpenAI-compatible endpoint at api.deepseek.com, and both Flash and Pro work through either. I migrated a test script in four minutes. Less dramatic than the model renaming suggests (DeepSeek Anthropic API guide).
Which DeepSeek models support the Responses API?
The support matrix was lopsided when I checked: only deepseek-v4-flash supported the Responses API, with deepseek-v4-pro support scheduled for early August 2026. DeepSeek added the route for Codex integrations (DeepSeek Responses API guide).
How does DeepSeek deduct API usage fees?
The formula is boring: tokens multiplied by price, deducted from the topped-up or granted balance, granted balance first when both are available. That last detail comes from DeepSeek alone. I found no independent source documenting the balance order, so treat it as single-source (DeepSeek pricing docs).
Methodology
Every number on this page was fetched from DeepSeek’s live docs on 2026-08-05, at 20:14 CEST, then cross-checked against two to four independent trackers per fact. I don’t publish AI pricing pages built from numbers I cannot defend.
I extracted 22 atomic facts, while marketing sentences went in the bin. Each fact got one of three labels:
Confirmed. Two or more independent sources agree.
Conflicting. Sources disagree, so I quote both and explain the likeliest cause.
Single-source. Only the official page documents it, with a dated caveat attached.
Every Source, With Fetch Dates
DeepSeek Models & Pricing, fetched 2026-08-05 twice (primary)
DeepSeek Rate Limit & Isolation, fetched 2026-08-05
DeepSeek Anthropic API guide, fetched 2026-08-05
DeepSeek Responses API guide, fetched 2026-08-05
DeepSeek pricing-details-usd (now 404), fetched 2026-08-05
Reuters, DeepSeek news hub, fetched 2026-08-05
InfoWorld, DeepSeek’s steep V4-Pro price cut, fetched 2026-08-05
MoneyDJ via TTV, permanent price cut, fetched 2026-08-05
Coworker AI, DeepSeek API pricing with changelog, fetched 2026-08-05
Codersera, DeepSeek V4 complete guide, fetched 2026-08-05
CrowdListen, DeepSeek V4 research, fetched 2026-08-05
Zenn, V4 preview-to-GA transition, fetched 2026-08-05
MACGPU, V4 full release pricing and benchmarks, fetched 2026-08-05
CloudZero, DeepSeek pricing 2026, fetched 2026-08-05
BenchLM, DeepSeek API pricing, fetched 2026-08-05
Verdent, V4 pricing and migration guide, fetched 2026-08-05
DevTk, DeepSeek V4 Pro specification, fetched 2026-08-05
MorphLLM, DeepSeek API 2026, fetched 2026-08-05
Reapi, V4 1M context guide, fetched 2026-08-05
ai-tldr.dev, peak and off-peak pricing, fetched 2026-08-05
TheRouter, peak-hour pricing routing, fetched 2026-08-05
StackFutures, V4 surge pricing, fetched 2026-08-05
Apidog, V4 Flash Responses API and Codex, fetched 2026-08-05
About This Page
I’m Karo Zieminski, founder of Product with Attitude. I fetch official vendor pages, extract the load-bearing facts, chase the discrepancies, and document the result before AI engines settle on the wrong number.
I know. Weird hobby.
Subscribe to Product with Attitude for verified AI pricing snapshots, discrepancy logs, and builder-side critical AI literacy every week.
Keep Reading
Perplexity Computer Pricing





