Flash 0731 vs Sonnet 5
Close intelligence, wild cost gap
Sonnet 5 is slightly ahead on the Artificial Analysis Intelligence Index (55 vs 52); DeepSeek V4 Flash wins on price, speed and open weights. Both offer a 1M context window.
Highlights
AA Intelligence
Sonnet 5 Adaptive Max vs Flash 0731 Max Effort. Overlap territory, not a blowout.
AA cost / task
Independent blended cost. Flash cache hits are extremely cheap.
Output tok/s
Flash is the faster decoder. Sonnet TTFT on max thinking can look huge (AA ~191s) because thinking counts.
Weights
Flash 0731 is MIT on HF. Sonnet 5 is proprietary, default in Claude Code.
The matchup
Companies ask if they can swap Sonnet 5 for Flash to cut bills. The honest answer is: for high-volume, well-specified coding/agent loops, often yes; for Claude-native tool fidelity and enterprise Claude stack, keep Sonnet.
What people argue
Community threads describe Sonnet 5 as a token guzzler while Flash cache hits cost almost nothing — consistent with the per-task gap Artificial Analysis measures ($0.06 vs $1.54). OpenCode's published volume charts also put Flash at the top of Go traffic.
Dates that matter
Sonnet 5 ships. Intro $2/$10 through 2026-08-31, then $3/$15.
Flash 0731 public beta.
DeepSeek peak/off-peak. Still far under Sonnet.
Read the table honestly
Use these numbers
AA compare page: 52 vs 55, $0.06 vs $1.54, 115 vs 69 tok/s, TTFT 1.43s vs 191.38s, both 1M context.
Conclusions this data does not support
That Flash “ranks the same as Sonnet 5 on everything.” DeepSWE anecdotes ≠ AA index. Official DeepSeek tables compare Flash to Opus-4.8 / GLM-5.2, not Sonnet 5.
When to pick whom
Pick Flash 0731
Volume, local/open, OpenCode Go, budget agent loops.
Pick Sonnet 5
Claude Code default, Bedrock/enterprise, instruction-fidelity, adaptive thinking productized.
How to use both
Same 1M window. Different tokenizer and thinking bills. Sonnet intro price ends 2026-09-01.
On QCode
QCode currently lists both deepseek-v4-flash and claude-sonnet-5 on /models, so you can compare them with one key. The live list on that page determines availability.
FAQ
Who is more intelligent?
Sonnet 5, slightly, on AA 55 vs 52.
Who is cheaper?
Flash, by a wide margin, even after 8-16 peak rates.
Why is Sonnet TTFT 191s?
Max adaptive thinking time is included. Not a hang.
Can Flash replace Claude Code?
It can be the model inside another harness (OpenCode, dsh). It is not the Claude Code product.
Sonnet price after Sep 1?
Standard $3 input / $15 output per million tokens.
Same context?
Both advertise 1M. Output caps differ (Flash 384K vs Sonnet 128K, 300K batches beta).
Sources
artificialanalysis.ai Flash vs Sonnet 5; DeepSeek pricing; Anthropic Sonnet 5 intro window.
Run both, route by job
One key should reach the cheap open lane and the Claude default — not a religious pick.
Related
Flash 0731 guide
Full 0731 specs.
Pro 0813
When Flash is not enough.
Sonnet 5 on this site
Existing Claude Sonnet 5 cluster.
Not affiliated with DeepSeek or Anthropic.