Flash 0731 vs Sonnet 5
Close intelligence, wild cost gap
Sonnet 5 is slightly ahead on the Artificial Analysis Intelligence Index (55 vs 52); DeepSeek V4 Flash wins on price, speed and open weights. Both offer a 1M context window.
Updated 2026-08-30
Highlights
AA Intelligence
Sonnet 5 Adaptive Max vs Flash 0731 Max Effort. Overlap territory, not a blowout.
AA cost / task
Independent blended cost. Flash cache hits are extremely cheap.
Output tok/s
Flash is the faster decoder. Sonnet TTFT on max thinking can look huge (AA ~191s) because thinking counts.
Weights
Flash 0731 is MIT on HF. Sonnet 5 is proprietary, default in Claude Code.
The matchup
Companies ask if they can swap Sonnet 5 for Flash to cut bills. The honest answer is: for high-volume, well-specified coding/agent loops, often yes; for Claude-native tool fidelity and enterprise Claude stack, keep Sonnet.
What people argue
Community threads describe Sonnet 5 as a token guzzler while Flash cache hits cost almost nothing — consistent with the per-task gap Artificial Analysis measures ($0.06 vs $1.54). OpenCode's published volume charts also put Flash at the top of Go traffic.
Dates that matter
Sonnet 5 ships at intro $2/$10. On 2026-08-10 that rate was made standard; the 2026-09-01 step to $3/$15 was cancelled.
Flash 0731 public beta.
DeepSeek peak/off-peak. Still far under Sonnet.
Read the table honestly
Use these numbers
AA compare page: 52 vs 55, $0.06 vs $1.54, 115 vs 69 tok/s, TTFT 1.43s vs 191.38s, both 1M context.
Conclusions this data does not support
That Flash “ranks the same as Sonnet 5 on everything.” DeepSWE anecdotes ≠ AA index. Official DeepSeek tables compare Flash to Opus-4.8 / GLM-5.2, not Sonnet 5.
When to pick whom
Pick Flash 0731
Volume, local/open, OpenCode Go, budget agent loops.
Pick Sonnet 5
Claude Code default, Bedrock/enterprise, instruction-fidelity, adaptive thinking productized.
How to use both
Same 1M window. Different tokenizer and thinking bills. Sonnet is $2/$10 as the standard rate (confirmed 2026-08-10), not $3/$15 after 2026-09-01.
On QCode
QCode currently lists both deepseek-v4-flash and claude-sonnet-5 on /models, so you can compare them with one key. The live list on that page determines availability.
FAQ
Who is more intelligent?
Sonnet 5, slightly, on AA 55 vs 52.
Who is cheaper?
Flash, by a wide margin, even after 8-16 peak rates.
Why is Sonnet TTFT 191s?
Max adaptive thinking time is included. Not a hang.
Can Flash replace Claude Code?
It can be the model inside another harness (OpenCode, dsh). It is not the Claude Code product.
Does Sonnet go to $3/$15 after Sep 1?
No. $2/$10 was confirmed as the standard rate on 2026-08-10; the Sep 1 hike will not happen.
Same context?
Both advertise 1M. Output caps differ (Flash 384K vs Sonnet 128K, 300K batches beta).
Sources
artificialanalysis.ai Flash vs Sonnet 5; DeepSeek pricing; Anthropic Sonnet 5 intro window.
Run both, route by job
One key should reach the cheap open lane and the Claude default — not a religious pick.
Related
Flash 0731 guide
Full 0731 specs.
Pro 0813
When Flash is not enough.
Sonnet 5 on this site
Existing Claude Sonnet 5 cluster.
Not affiliated with DeepSeek or Anthropic.