Official pricing · 2026-09-18

DeepSeek V4.1 Flash pricing: peak/off-peak, cache-hit rate, how the bill adds up
Peak, cache hit and how the bill adds up

DeepSeek prices V4.1 Flash per 1M tokens, with off-peak exactly half of peak; 1M context and 384K max output. All six rates and the peak windows, converted to your local timezone, are below.

Updated 2026-09-19

#Peak/off-peak#Cache hit $0.003#1M context#384K output

Six rates and two limits

$0.15 / $0.3

Input (cache miss)

$0.15 off-peak, $0.3 peak, per 1M tokens.

$0.6 / $1.2

Output

$0.6 off-peak, $1.2 peak, per 1M tokens.

$0.003 / $0.006

Input (cache hit)

Cache-hit input drops to $0.003 off-peak / $0.006 peak.

2500

Concurrency

Flash 2500; V4 Pro on the same page is 500.

What this is

The official sheet quotes per 1M tokens, splits input into cache-hit and cache-miss, then into off-peak and peak. Off-peak is half of peak; peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, everything else off-peak.

What happened

New pricing took effect 04:00 UTC on 2026-09-10, when V4.1 Flash shipped; the changelog's 2026-09-10 entry confirmed V4 Pro continues with unchanged billing.

Timeline

2026-08-13

2026-08-13 DeepSeek-V4-Pro-0813 went GA — the V4 Pro column on today's sheet.

2026-09-10

2026-09-10 04:00 UTC V4.1 Flash live with the new pricing (from $0.15 / $0.6).

2026-09-14

The changelog's 2026-09-10 entry and the pricing footnote confirm V4 Pro continues, billing unchanged.

Confirmed vs caution

Officially confirmed

The pricing page names the live versions DeepSeek-V4.1-Flash and DeepSeek-V4-Pro-0813 (GA 2026-08-13), and states off-peak is half of peak with peak hours 01:00-04:00 and 06:00-10:00 UTC weekdays.

⚠️ Caution

No end date is published; the footnote says prices may change and the page is authoritative. These three rates are the ones in force since 2026-09-10 04:00 UTC — don't lock them into a long-term customer quote.

Same vs differs

Same

Both priced per 1M tokens, three rates (cache-hit input, cache-miss input, output), both split peak/off-peak, off-peak half of peak.

Differs

Cache-miss input $0.66 vs Flash $0.15; output $1.98 vs $0.6; concurrency 500 vs 2500; V4 Pro has no vision.

What to do

Apply the three rates to your own input/output mix, then decide whether to shape long context as a cache-hittable prefix. Repeating the same prompt drops input cost by an order of magnitude.

On QCode

On QCode set deepseek-v4.1-flash; billed at official rate × service fee, same peak and cache semantics. To compare, switch to deepseek-v4-pro ($0.66 / $1.98) with the same key.

FAQ

What are the six rates?

Cache-miss input $0.15 off-peak / $0.3 peak; output $0.6 / $1.2; cache-hit input $0.003 / $0.006 — all per 1M tokens.

When is peak?

Peak is 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; all other hours are off-peak, weekends included. That is 09:00-12:00 and 14:00-18:00 Beijing, 10:00-13:00 and 15:00-19:00 Tokyo, 04:00-07:00 and 09:00-13:00 Moscow.

How much cheaper than V4 Pro?

Cache-miss input $0.15 vs $0.66, output $0.6 vs $1.98, cache-hit input $0.003 vs $0.022. Flash allows 2500 concurrent requests against 500, and V4 Pro has no vision.

When does cache hit pay off?

A hit takes input from $0.15 to $0.003 — 1/50. Put the unchanging system prompt and long documents at the front and vary only the tail.

Does thinking mode cost more?

Thinking emits extra thinking tokens billed as output. At the same $1.2/M peak output rate, a longer chain of thought means a bigger bill — set a non-thinking baseline first.

Will this price hold?

The footnote reserves the right to adjust and points back to the official page, which is why this page carries a capture date of 2026-09-18. Re-check before quoting long term.

Sources

DeepSeek official Models & Pricing, changelog and news260910, captured 2026-09-18. All rates per 1M tokens.

Run your own bill at official rate × fee

One QCode key: set deepseek-v4.1-flash, switch to deepseek-v4-pro to compare.

Related

This page restates the official rate table and windows only — no third-party benchmarks, no speculation. DeepSeek reserves the right to adjust pricing.