Comparison · live models

Flash 0731 vs Sonnet 5
Close intelligence, wild cost gap

Sonnet 5 is slightly ahead on the Artificial Analysis Intelligence Index (55 vs 52); DeepSeek V4 Flash wins on price, speed and open weights. Both offer a 1M context window.

#AA 52 vs 55#$0.06 vs $1.54#115 vs 69 t/s#1M vs 1M

Highlights

52 vs 55

AA Intelligence

Sonnet 5 Adaptive Max vs Flash 0731 Max Effort. Overlap territory, not a blowout.

$0.06 vs $1.54

AA cost / task

Independent blended cost. Flash cache hits are extremely cheap.

115 vs 69

Output tok/s

Flash is the faster decoder. Sonnet TTFT on max thinking can look huge (AA ~191s) because thinking counts.

Open vs closed

Weights

Flash 0731 is MIT on HF. Sonnet 5 is proprietary, default in Claude Code.

The matchup

Companies ask if they can swap Sonnet 5 for Flash to cut bills. The honest answer is: for high-volume, well-specified coding/agent loops, often yes; for Claude-native tool fidelity and enterprise Claude stack, keep Sonnet.

What people argue

Community threads describe Sonnet 5 as a token guzzler while Flash cache hits cost almost nothing — consistent with the per-task gap Artificial Analysis measures ($0.06 vs $1.54). OpenCode's published volume charts also put Flash at the top of Go traffic.

Dates that matter

2026-06-30

Sonnet 5 ships. Intro $2/$10 through 2026-08-31, then $3/$15.

2026-07-31

Flash 0731 public beta.

2026-08-16

DeepSeek peak/off-peak. Still far under Sonnet.

Read the table honestly

Use these numbers

AA compare page: 52 vs 55, $0.06 vs $1.54, 115 vs 69 tok/s, TTFT 1.43s vs 191.38s, both 1M context.

Conclusions this data does not support

That Flash “ranks the same as Sonnet 5 on everything.” DeepSWE anecdotes ≠ AA index. Official DeepSeek tables compare Flash to Opus-4.8 / GLM-5.2, not Sonnet 5.

When to pick whom

Pick Flash 0731

Volume, local/open, OpenCode Go, budget agent loops.

Pick Sonnet 5

Claude Code default, Bedrock/enterprise, instruction-fidelity, adaptive thinking productized.

How to use both

Same 1M window. Different tokenizer and thinking bills. Sonnet intro price ends 2026-09-01.

On QCode

QCode currently lists both deepseek-v4-flash and claude-sonnet-5 on /models, so you can compare them with one key. The live list on that page determines availability.

FAQ

Who is more intelligent?

Sonnet 5, slightly, on AA 55 vs 52.

Who is cheaper?

Flash, by a wide margin, even after 8-16 peak rates.

Why is Sonnet TTFT 191s?

Max adaptive thinking time is included. Not a hang.

Can Flash replace Claude Code?

It can be the model inside another harness (OpenCode, dsh). It is not the Claude Code product.

Sonnet price after Sep 1?

Standard $3 input / $15 output per million tokens.

Same context?

Both advertise 1M. Output caps differ (Flash 384K vs Sonnet 128K, 300K batches beta).

Sources

artificialanalysis.ai Flash vs Sonnet 5; DeepSeek pricing; Anthropic Sonnet 5 intro window.

Run both, route by job

One key should reach the cheap open lane and the Claude default — not a religious pick.