Historical record · 2026-07-31 build

DeepSeek V4 Flash 0731
Historical record: DeepSeek retired this version on 2026-09-10

Public beta API on 2026-07-31. Model id stays deepseek-v4-flash. Agent benches jumped; V4-Pro and the app were unchanged that day.

Updated 2026-09-20

#deepseek-v4-flash#0731#TB2.1 82.7#$0.14 / $0.28

Highlights

deepseek-v4-flash

Same API name

Point the existing Flash id at 0731. No client rename.

1M / 384K

Context / max out

Official pricing page. Thinking effort: low / high / max.

$0.14 / $0.28

Cache miss / output

Cache hit $0.0028. Peak/off-peak from 2026-08-16 16:00 UTC.

82.7

Terminal Bench 2.1

Official: Harness Minimal + max. DeepSWE 54.4 (Preview was 7.3).

What 0731 actually is

This page records what DeepSeek officially said about the 2026-07-31 Flash post-training at the time: same Flash structure and scale, re-done post-training; HF lists about 304B (community shorthand often cites 284B/13B active MoE); MIT weights; the checkpoint ships DSpark speculative decoding. On 2026-09-10 DeepSeek released V4.1-Flash and retired the deepseek-v4-flash api name — those requests are served by DeepSeek-V4.1-Flash and billed at the Flash price. So this is a historical comparison, not the current version.

Why the circle flipped overnight

Reporting and official benchmarks at the time (2026-07-31): the preview Flash felt mediocre, 0731's AA approached the March 2026 frontier (about 50–52) and made it OpenCode's volume leader; LocalLLaMA called it "five-month-old frontier you can run at home". All of that is the 0731 narrative, not the current availability — the live Flash tier is V4.1-Flash from 2026-09-10.

Timeline

2026-04-24

2026-04-24 V4-Pro / V4-Flash entered the API (preview).

2026-07-31

2026-07-31 public beta of Flash 0731. Changelog: Flash only, naming the Responses API and Codex.

2026-09-10

2026-09-10 DeepSeek released V4.1-Flash; deepseek-v4-flash was retired and routed to DeepSeek-V4.1-Flash, billed at the Flash price (pricing footnote (1)).

Confirmed facts vs claims to treat carefully

Confirmed

The benchmarks above are 0731's own claims at the time; 1M/384K; MIT; three effort levels; native Responses API; evaluation ran in DeepSeek Harness Minimal. DeepSeek's 2026-09-10 release note and pricing footnote confirm deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired, served by DeepSeek-V4.1-Flash and billed at the Flash price.

Claims worth treating with caution

AA 52 against Sonnet 5's 55 is an independent evaluation, not an official "we beat Sonnet". The "196GB" in the blog is not a spec sheet. And "0731 is still today's Flash" is not true either: DeepSeek retired it on 2026-09-10.

Quick compare

vs Claude Sonnet 5

AA 52 vs 55. Cost/task ~$0.06 vs $1.54. Flash is faster and open. Sonnet wins instruction stack and Claude Code defaults.

vs V4-Pro 0813

0731 beat Preview Pro. After 8-13, 0813 is the flagship (TB2.1 87.9, DeepSWE 62.7). Flash stays the default cheap lane.

How to call it

Official: api.deepseek.com, model deepseek-v4-flash. Also Anthropic-compatible path and Responses/Codex docs. Local: HF deepseek-ai/DeepSeek-V4-Flash-0731 with vLLM/SGLang DSpark flags. OpenCode Go/Free also list Flash.

On QCode

Do not build against the 0731 version: DeepSeek retired the deepseek-v4-flash name, and requests with it are served by DeepSeek-V4.1-Flash at the Flash price. For new integrations set deepseek-v4.1-flash as listed in QCode's /models; for the Pro tier use deepseek-v4-pro. This page makes no availability claim for 0731.

Flash 0731 FAQ

Is 0731 a new model?

Same Flash structure with new post-training, and the API name was unchanged for the 2026-07-31 release. But since 2026-09-10 DeepSeek retired deepseek-v4-flash and routes it to V4.1-Flash, so this only describes 0731 at the time.

Can I run it locally?

Yes, MIT weights. Full size needs serious unified memory or multi-GPU; 16GB laptops want a quant, not the full checkpoint.

Is it better than Sonnet 5?

On AA intelligence, slightly behind (52 vs 55). On price and speed, far ahead. Pick by stack, not a trophy.

Does 8-16 pricing kill the deal?

Peak output becomes $1.32; off-peak $0.66. Still well under Sonnet 5’s $10/$15 band.

Why are official benches so high vs my Claude Code run?

They used Harness Minimal + max. A different harness changes the score.

Flash or Pro?

Default Flash (concurrency 2500). Escalate hard agent jobs to Pro 0813 (concurrency 500).

Sources

DeepSeek changelog, the news260910 release note and Models & Pricing (captured 2026-09-18); the 0731-era official blog and third-party evaluations. Retirement and temporary routing rest on pricing footnote (1).

Put Flash in a multi-model mix

This page is the 0731 historical record. For new integrations set deepseek-v4.1-flash from QCode's /models, or deepseek-v4-pro for the Pro tier.

Related

This page keeps the 2026-07-31 official benchmarks and third-party numbers from that time without claiming the version is still available. The live Flash tier is V4.1-Flash from 2026-09-10; defer to DeepSeek's docs and QCode's /models.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.