Official Flash · 2026-07-31

DeepSeek V4 Flash 0731
Same size, new post-train

Public beta API on 2026-07-31. Model id stays deepseek-v4-flash. Agent benches jumped; V4-Pro and the app were unchanged that day.

#deepseek-v4-flash#0731#TB2.1 82.7#$0.14 / $0.28

Highlights

deepseek-v4-flash

Same API name

Point the existing Flash id at 0731. No client rename.

1M / 384K

Context / max out

Official pricing page. Thinking effort: low / high / max.

$0.14 / $0.28

Cache miss / output

Cache hit $0.0028. Peak/off-peak from 2026-08-16 16:00 UTC.

82.7

Terminal Bench 2.1

Official: Harness Minimal + max. DeepSWE 54.4 (Preview was 7.3).

What 0731 actually is

DeepSeek says 0731 keeps the Flash architecture and size and only re-does post-training. Hugging Face lists ~304B params (community often says 284B/13B active MoE). Weights are MIT. DSpark speculative decoding ships in the checkpoint.

Why the circle flipped overnight

Preview Flash felt mid. 0731 landed near March-2026 frontier on AA (~50–52) and became OpenCode’s volume king. LocalLLaMA treated it as “frontier from five months ago, runnable at home.” That is the SEO story: cheap, open, suddenly agentic.

Timeline

2026-04-24

V4-Pro and V4-Flash enter the API (preview).

2026-07-31

Flash 0731 public beta. Changelog: only Flash upgrades; Pro/app unchanged. Responses API + Codex notes.

2026-08-16 16:00 UTC

Peak/off-peak billing starts (off-peak = half of peak).

Confirmed facts vs claims to treat carefully

Confirmed

Official benches above; 1M/384K; MIT weights; three effort levels; native Responses API. Evaluated in DeepSeek Harness Minimal.

Claims worth treating with caution

AA 52 vs Sonnet 5’s 55 is independent, not an official “we beat Sonnet” line. Hardware floors like “196GB” are blogger notes, not a datasheet.

Quick compare

vs Claude Sonnet 5

AA 52 vs 55. Cost/task ~$0.06 vs $1.54. Flash is faster and open. Sonnet wins instruction stack and Claude Code defaults.

vs V4-Pro 0813

0731 beat Preview Pro. After 8-13, 0813 is the flagship (TB2.1 87.9, DeepSWE 62.7). Flash stays the default cheap lane.

How to call it

Official: api.deepseek.com, model deepseek-v4-flash. Also Anthropic-compatible path and Responses/Codex docs. Local: HF deepseek-ai/DeepSeek-V4-Flash-0731 with vLLM/SGLang DSpark flags. OpenCode Go/Free also list Flash.

On QCode

QCode currently lists deepseek-v4-flash on /models. That live list is the source of truth for what you can call, and the same API key also covers Claude, GPT, Gemini and other families.

Flash 0731 FAQ

Is 0731 a new model?

Same Flash structure, new post-train. API name unchanged.

Can I run it locally?

Yes, MIT weights. Full size needs serious unified memory or multi-GPU; 16GB laptops want a quant, not the full checkpoint.

Is it better than Sonnet 5?

On AA intelligence, slightly behind (52 vs 55). On price and speed, far ahead. Pick by stack, not a trophy.

Does 8-16 pricing kill the deal?

Peak output becomes $1.32; off-peak $0.66. Still well under Sonnet 5’s $10/$15 band.

Why are official benches so high vs my Claude Code run?

They used Harness Minimal + max. A different harness changes the score.

Flash or Pro?

Default Flash (concurrency 2500). Escalate hard agent jobs to Pro 0813 (concurrency 500).

Sources

Sources: the 2026-07-31 update note and pricing page on api-docs.deepseek.com, the DeepSeek-V4-Flash-0731 model card on Hugging Face, and Artificial Analysis' Flash vs Sonnet 5 comparison. Pricing moves to peak/off-peak rates from 2026-08-16 16:00 UTC.

Put Flash in a multi-model mix

Use 0731 where cost and volume win; keep Claude / GPT for the jobs that still need them.