Flash-tier comparison · both on the list

GLM-5.3-Flash vs DeepSeek V4 Flash
List $0.15/$0.50 vs $0.14/$0.28; both 1M windows, MIT weights

Zhipu shipped GLM-5.3-Flash on 2026-08-26 (320B-A18B, 1M, MIT). Official list $0.15 / $0.50 per million; launch half-off $0.075 / $0.25 until 2026-09-09 24:00 UTC+8. DeepSeek V4 Flash 0731 lists uncached / output $0.14 / $0.28, window 1M / 384K. Both ids have usage > 0 on this site in 30 days — one key, one endpoint, switch the model field.

#glm-5.3-flash#deepseek-v4-flash#$0.15/$0.50 vs $0.14/$0.28#both 1M

Four numbers you can quote

$0.15 / $0.50

GLM-5.3-Flash official list

In / out per million. Launch half-off is $0.075 / $0.25 through 2026-09-09 24:00 UTC+8. QCode follows official × fee; on 2026-08-27 the billed pair was $0.14 / $0.49.

$0.14 / $0.28

DeepSeek V4 Flash uncached / output

0731 official rates. Peak/off-peak started 2026-08-16 16:00 UTC. Window 1M in / 384K max out.

1M / 1M

Both pitch a million-token context

GLM-5.3-Flash official 1M. DeepSeek V4 Flash 1M / 384K. Long-repo jobs: watch the real cut, not only the marketing window.

MIT / MIT

Both ship open weights

GLM-5.3-Flash landed on HF 2026-08-26. DeepSeek V4 Flash-0731 is MIT too. Open weights ≠ you must self-host; the API path follows official list prices.

Flash versus Flash, not another GLM guide

The site already has a GLM-5.3-Flash guide, a DeepSeek V4 Flash 0731 page, and GLM-5.2 vs DeepSeek V4 Pro. This page only lines up the cheap tiers: official price, window, license. GLM is a 320B-A18B multimodal MoE. DeepSeek 0731 is the same Flash scale with fresh post-training and a jumped agent score. Pick by task, not slogan.

Why Flash-vs-Flash showed up after 08-26

Launch day ended the Ox Alpha stealth trial and swapped in z-ai/glm-5.3-flash. DeepSeek Flash 0731 had been in preview since 07-31, with peak/off-peak in August. Official output is about $0.50 vs $0.28; input sits almost on top of each other. After QCode added GLM-5.3-Flash on 2026-08-27, both ids take real traffic on one endpoint.

Timeline

2026-07-31

DeepSeek V4 Flash 0731 API preview. Id stays deepseek-v4-flash. Official uncached / output $0.14 / $0.28.

2026-08-16 16:00 UTC

DeepSeek Flash moves to peak/off-peak. Do not quote only the daytime list.

2026-08-26 / 2026-08-27

GLM-5.3-Flash ships and opens weights; QCode adds it the next day at $0.14 / $0.49 (official × fee). Official half-off runs through 2026-09-09 24:00 UTC+8.

Confirmed vs misread

Confirmed

Both official prices, 1M windows, MIT, GLM ship date 2026-08-26, discount end 2026-09-09 24:00 UTC+8, DeepSeek peak/off-peak from 08-16, both ids usage > 0 here — vendor pages and already-shipped on-site copy, rechecked 2026-08-30.

Misread

“Flash means a laptop runs it” — GLM is 320B total; that is not a phone. “Output prices match” — official $0.50 vs $0.28. “QCode equals official half-off” — QCode is official × fee; do not merge 2026-08-27’s $0.14/$0.49 with $0.075/$0.25.

When to pick which

Lean GLM-5.3-Flash

Need vision in, the Zhipu tool chain, or the official half-off before 09-09. Architecture is 320B-A18B multimodal. Launch cache-read list $0.015. Trust the console for the bill.

Lean DeepSeek V4 Flash

Want the lower official output rate, peak/off-peak, or you already sit on 0731’s agent jump. Terminal Bench 2.1 public 82.7 lives on the 0731 page. Check 384K max out against the job.

How to switch on one endpoint (both sold)

The ids are glm-5.3-flash and deepseek-v4-flash. One already-provisioned key, one Chat Completions endpoint, change model. Estimate on official lists, then read the fee-multiplied number on your account. Do not treat the half-off window as a forever price.

Both Flash ids are callable on QCode

glm-5.3-flash and deepseek-v4-flash both show usage > 0 in the 30-day table. Prepaid balance at official price × service fee. GLM on 2026-08-27 listed $0.14/M in, $0.49/M out, $0.14/M cache read. DeepSeek follows $0.14/$0.28 and peak/off-peak. Same endpoint; no second console deposit.

FAQ

Which table is the official price?

GLM: Zhipu / Z.ai $0.15/$0.50, half-off $0.075/$0.25 until 2026-09-09 24:00 UTC+8. DeepSeek: uncached / output $0.14/$0.28 plus peak/off-peak. QCode bill = official × fee.

Are both windows 1M?

GLM official 1M. DeepSeek Flash 1M in, 384K max out. Long-output jobs: check 384K first.

How is this different from GLM-5.2 vs DeepSeek V4 Pro?

That page is flagship / Pro. This page is Flash only. Do not copy Pro scores onto the Flash row.

Self-host pick?

Both MIT. GLM 320B-A18B VRAM is not the word “Flash.” DeepSeek 0731 has its own homelab threads. This page is API selection, not a rack list.

Can I try both on QCode?

Yes. Both ids have real usage. Same key, change model. Prices follow upstream.

Is GLM still half-off after 09-09?

The official discount is written through 2026-09-09 24:00 UTC+8. After that, list $0.15/$0.50 unless Zhipu posts otherwise. Do not freeze the discount as permanent.

Sources

GLM list, half-off window, 320B-A18B, 1M, MIT: Zhipu / Z.ai 2026-08-26 plus on-site glm-5-3-flash guide and pricing page (rechecked 2026-08-30). QCode billed $0.14/$0.49: that page’s qcode_text, 2026-08-27. DeepSeek $0.14/$0.28, 1M/384K, peak/off-peak 2026-08-16: on-site deepseek-v4-flash-0731 and api-docs.deepseek.com. Usage: this session’s 30-day table, both ids > 0.

Both Flash models are sold; change model on the same endpoint

Official price × service fee. GLM half-off has an end date; DeepSeek has peak/off-peak. Estimate, then shift traffic.

Related

Prices and windows follow Zhipu / DeepSeek official pages and your bill. Discounts and peak/off-peak move. QCode bills official × fee and does not promise to match official half-off to the cent.