DeepSeek V4 Price Hike, Fully Explained
After peak/off-peak pricing landed
Starting 08-17 the DeepSeek API moved to peak/off-peak pricing: V4-Pro peak output ¥27 per million tokens (+350%), half price off-peak. If you can schedule, your costs may actually go down.
Updated 2026-09-20
Four key numbers
V4-Pro peak output increase
Peak output went from ¥6 to ¥27 per million tokens. Cache-hit input rose the most — up to 1100%.
V4-Flash max increase
The lightweight tier went up in step (per Caixin, 08-21). Off-peak remains half of peak.
Peak windows (Beijing time)
Weekdays 9:00-12:00 and 14:00-18:00 are peak; 00:30-08:30, weekends and public holidays run at off-peak rates.
Effective date
Console warning on 08-06 → official announcement on 08-13 (same day as V4 Pro GA) → effective at midnight 08-17.
What actually changed
DeepSeek moved its API from flat pricing to peak/off-peak time-of-use pricing: every billable item goes up during peak hours and returns to half price off-peak. The official framing is 'a return to value-based token pricing' — the 2.5x promotional discount from May and the low-price strategy are over, freeing compute budget to 'develop stronger large models.' The hike was announced the same day as V4 Pro GA (0813) and took effect within two weeks.
Market reaction
The price-rise note landed alongside “V4 Pro GA with much stronger agent capability”, read across the industry as DeepSeek moving from “price slayer” to value pricing. The table was rewritten twice after that: on 2026-09-10 new rates came in with V4.1-Flash and the Flash tier was cut, and the changelog's 2026-09-10 entry withdrew the V4-Pro redirect, keeping V4 Pro with billing unchanged. The developer conversation is still about re-running the numbers: off-peak scheduling plus cache-hit optimisation.
Timeline
Developer console warning: 'we plan a broad API price increase soon, expected to be significant.'
Official announcement: peak/off-peak pricing details + V4 Pro GA on the same day.
2026-08-17 the new rates take effect. On 2026-08-21 V4-Flash-Vision-Exp shipped at the V4-Flash price; on 2026-09-10 both it and deepseek-v4-flash were retired, with requests served by V4.1-Flash at the Flash price.
Confirmed vs watch out
Confirmed
The peak/off-peak windows, V4-Pro peak output at ¥27 per million tokens, the +350%/+1100%/+400% increase figures, and the 08-17 effective date all appear in the official announcement and in Caixin and guancha.cn reporting.
Watch out
The '11x price increase' headline refers to a single item — cache-hit input (¥0.025→¥0.3). Not everything went up 11x. Do the math on your own call profile; don't let the single largest item set the narrative.
How to cut costs after the hike
Off-peak scheduling
Move batch jobs, evals and nightly builds into off-peak windows (Beijing 00:30-08:30 plus weekends) and the price is cut in half outright.
Caching and model tiering
Reuse long prefixes to harvest the cache-hit discount; downgrade chore work to Flash or another cheap model, and reserve Pro for the hard stuff.
How to run your numbers
Three steps: ① pull your call profile (input/output/cache-hit ratios); ② re-price it under the new peak and off-peak rates; ③ push every deferrable workload into off-peak windows. For most batch-heavy workloads the adjusted total is flat or even lower.
On QCode
QCode's list tracks the official price (official × service fee). The two live tiers are deepseek-v4.1-flash (official off-peak $0.15 / $0.6) and deepseek-v4-pro (from $0.66 / $1.98). Since 2026-09-10 deepseek-v4-flash is retired: requests on that name are served by V4.1-Flash at the Flash price. Peak/off-peak differences follow the console's live numbers; off-peak calls get the lower rate automatically.
FAQ
When did DeepSeek raise prices?
Effective at midnight (Beijing time) on 2026-08-17. The console warning came on 08-06 and the official announcement on 08-13, the same day as V4 Pro GA.
How much did prices rise?
For 2026-08-17: V4-Pro peak output ¥27 per million tokens (about +350%), cache-hit input up to +1100%, V4-Flash up to +400%, with off-peak at half the peak. From 2026-09-10 the Flash tier came down again, and the live id is deepseek-v4.1-flash.
How are peak and off-peak windows defined?
Peak: weekdays 9:00-12:00 and 14:00-18:00 Beijing time. Off-peak: 00:30-08:30, weekends and statutory holidays.
Is DeepSeek still worth it after the hike?
Off-peak pricing is still among the cheapest of any open-source flagship. If you can schedule and use caching, the real-world cost increase is manageable.
Do QCode prices go up too?
QCode passes through the official price times the service rate; the /models catalog is live. Off-peak calls are billed at off-peak rates.
Is the '11x increase' real?
The single largest item is real (cache-hit input ¥0.025→¥0.3), but most billable items rose far less. Price your own call profile, not the biggest single number.
Sources
DeepSeek's pricing note (2026-08-13, reported by Xinhua Finance), Guancha (2026-08-17), Caixin (2026-08-21), Kaiyuan Securities (2026-08-06), Xueqiu summaries; plus the news260910 note of 2026-09-10, the changelog's 2026-09-10 entry and pricing footnote (1) (archived 2026-09-18).
Run batches off-peak, halve the cost
QCode passes through official pricing times the service rate — off-peak calls automatically get the lower rate.
Related reading
DeepSeek V4 Pro 0813
The GA build's specs and how to call it.
DeepSeek V5 tracker
The 'stronger model' rumors behind the price hike.
QCode pricing guide
Service rates and billing terms.
Not affiliated with DeepSeek. Prices follow the official announcement and QCode's live /models listing.