Comparison · live models

Flash 0731 vs Sonnet 5
Close intelligence, wild cost gap (0731 is a historical build)

Sonnet 5 is slightly ahead on the Artificial Analysis Intelligence Index (55 vs 52); DeepSeek V4 Flash wins on price, speed and open weights. Both offer a 1M context window.

Updated 2026-09-19

#AA 52 vs 55#$0.06 vs $1.54#115 vs 69 t/s#1M vs 1M

Highlights

52 vs 55

AA Intelligence

Sonnet 5 Adaptive Max vs Flash 0731 Max (the 2026-07-31 build). Overlap territory, not a blowout.

$0.06 vs $1.54

AA cost / task

Independent blended cost. Flash cache hits are extremely cheap.

115 vs 69

Output tok/s

Flash is the faster decoder. Sonnet TTFT on max thinking can look huge (AA ~191s) because thinking counts.

Open vs closed

Weights

Flash 0731 ships MIT weights on HF; since 2026-09-10 requests to that id are answered by deepseek-v4.1-flash. Sonnet 5 is proprietary, default in Claude Code.

The matchup

Companies ask if they can swap Sonnet 5 for Flash to cut bills. The honest answer is: for high-volume, well-specified coding/agent loops, often yes; if what you need is fidelity to Claude's own tools and the enterprise Claude stack, keep Sonnet.

What people argue

Community threads describe Sonnet 5 as a token guzzler while Flash cache hits cost almost nothing — consistent with the per-task gap Artificial Analysis measures ($0.06 vs $1.54). OpenCode's published volume charts also put Flash at the top of Go traffic.

Dates that matter

2026-06-30

Sonnet 5 ships at intro $2/$10. On 2026-08-10 that rate was made standard; the step to $3/$15 planned for 2026-09-01 was cancelled.

2026-07-31

2026-07-31 public beta of Flash 0731 (this build was retired on 2026-09-10).

2026-08-16

2026-08-16 DeepSeek peak/off-peak pricing. Still far under Sonnet.

Read the table honestly

Use these numbers

AA compare page: 52 vs 55, $0.06 vs $1.54, 115 vs 69 tok/s, TTFT 1.43s vs 191.38s, both 1M context.

Conclusions this data does not support

That Flash “ranks the same as Sonnet 5 on everything.” DeepSWE anecdotes ≠ AA index. Official DeepSeek tables compare Flash to Opus-4.8 / GLM-5.2, not Sonnet 5.

When to pick whom

Pick Flash 0731 (historical)

Volume, local/open, OpenCode Go, budget agent loops.

Pick Sonnet 5

Claude Code default, Bedrock/enterprise, instruction-fidelity, adaptive thinking productized.

How to use both

Same 1M window. Different tokenizer and thinking bills. Sonnet is $2/$10 as the standard rate (confirmed 2026-08-10); the $3/$15 planned from 2026-09-01 will not happen.

On QCode

QCode's /models currently lists deepseek-v4.1-flash and claude-sonnet-5 — those are the two ids to compare today. The legacy name deepseek-v4-flash was retired by the vendor on 2026-09-10, so the 0731 figures belong to an earlier generation. One key covers both; the live list on that page determines availability.

FAQ

Who is more intelligent?

Sonnet 5, slightly, on AA 55 vs 52.

Who is cheaper?

Flash, by a wide margin, even after 8-16 peak rates.

Why is Sonnet TTFT 191s?

Max adaptive thinking time is included. Not a hang.

Can Flash replace Claude Code?

It can be the model inside another harness (OpenCode, dsh). It is not the Claude Code product.

Does Sonnet go to $3/$15 after Sep 1?

No. $2/$10 was confirmed as the standard rate on 2026-08-10; the 2026-09-01 step to $3/$15 will not happen.

Same context?

Both advertise 1M. Output caps differ (Flash 384K vs Sonnet 128K, 300K batches beta).

Sources

artificialanalysis.ai Flash vs Sonnet 5; DeepSeek pricing; Anthropic Sonnet 5 intro window.

Run both, route by job

One key should reach the cheap open lane and the Claude default — not a religious pick.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.