Price tracker · 50% off ends September 9

GLM-5.3-Flash pricing
$0.075/$0.25 until September 9

List price is $0.15 in / $0.50 out per million tokens, halved for the launch window through 2026-09-09 24:00 (UTC+8). The savings are real — but the cache-read line deserves its own calculation.

#$0.075 / $0.25#50% off to 9/9#cache $0.015#1/40 of Opus 4.8

Four price lines

$0.075

Input · promo

Per million tokens. List is $0.15; it reverts when the promo window closes.

$0.25

Output · promo

Per million tokens. List is $0.50. A 1:3.3 in/out ratio, gentler than most models in this tier.

$0.015

Cache read · promo

List is $0.03. ⚠️ This is the line that dominates real cost for long-context agents — see the comparison below.

1/40

vs Opus 4.8

Both sit at AA Intelligence Index 57; at promo rates this runs about a fortieth of Claude Opus 4.8.

How the official price works

Z.ai and bigmodel.cn list GLM-5.3-Flash at $0.15 per million input tokens, $0.50 per million output, and $0.03 for cache reads. The launch promotion halves all three — $0.075 / $0.25 / $0.015 — and cache storage is free for the duration. That puts it at roughly a tenth of the same-generation GLM-5.3 (a twentieth during the promo) and a fortieth of Claude Opus 4.8. The deadline is explicit: 9 September 2026 at 24:00 (UTC+8).

What the aggregators charge

OpenRouter's z-ai/glm-5.3-flash currently shows the promo rate of $0.075 / $0.25 with $0.015 cache reads. Novita lists zai-org/glm-5.3-flash on its own serverless at $0.15 / $0.50. AIHubMix launched with a 50% limited-time discount. QuickSilver Pro claims roughly 20% below the official promo ($0.06 / $0.20). The spread comes from each vendor's subsidy strategy and settlement terms — before switching, confirm they actually serve the same 1M context and the same cache billing.

Price timeline

2026-08-26

Launch and API opening, with the 50% promo live from day one: $0.075 / $0.25, cache read $0.015.

2026-08-26

OpenRouter retires the free stealth/ox-alpha slug and lists z-ai/glm-5.3-flash at promo rates.

2026-09-09

24:00 (UTC+8): the promo ends. Back to $0.15 / $0.50 with cache reads at $0.03.

Confirmed vs verify yourself

Confirmed

List pricing of $0.15/$0.50, the promo rate of $0.075/$0.25, cache reads at $0.03 ($0.015 during the promo), the 2026-09-09 24:00 (UTC+8) deadline, and free cache storage during the window — all from Z.ai and bigmodel.cn documentation and launch announcements.

Verify yourself

Aggregator pricing moves constantly and does not always include the same context ceiling or cache policy; when a platform undercuts the official promo, check whether that is a subsidy or a downgrade. Coding Plan quota conversion (Flash at roughly 3x GLM-5.3) also varies by tier — read the official page for your region.

Cheap on paper, but price the cache first

The headline rate really is low

$0.075 / $0.25 is close to the floor for anything at AA index 57. For one-shot calls, short contexts and batch jobs the saving lands directly.

Cache reads are a different story

Cache read is $0.03 ($0.015 during the promo). One developer measured that at roughly 4x DeepSeek V4 Flash's off-peak rate. Agent loops re-read the same long context over and over, so cache charges can exceed what the headline rate saved. Work it out against your own hit rate and average context length first.

Running your own numbers

Split a typical request into three parts: fresh input tokens, cached input tokens, and output tokens. During the promo the formula is (fresh × $0.075 + cached × $0.015 + output × $0.25) / 1,000,000. Agent workflows usually land between 60% and 90% cache hit rate, and the higher that goes, the more the cache line dominates — which is exactly why the headline rate alone is misleading. After September 9, double all three numbers for your long-run cost.

On QCode

The Z.ai models routable on QCode are GLM-5.3, GLM-5.2 and GLM-5.1, billed at official rates times a service multiplier and tracking upstream changes automatically. GLM-5.3-Flash is not connected — this page is pricing reference only. To compare real costs across Chinese models today, start with GLM-5.3 or the DeepSeek V4 family, which additionally has off-peak pricing at half rate.

FAQ

Exactly when does the discount end?

9 September 2026 at 24:00, UTC+8. After that input returns to $0.15, output to $0.50 and cache reads to $0.03, all per million tokens.

Is there a free tier?

There is no permanently free public endpoint. What was free was the anonymous stealth/ox-alpha preview, removed from OpenRouter's production catalog on reveal day. New accounts on Z.ai's platform receive trial credit, which is a trial rather than a free tier.

How is caching billed?

Cache reads are $0.03 per million tokens ($0.015 during the promo) and cache storage is free for the promo window. Note that although cache read is a fifth of the input rate, in high-hit-rate long-context work it accounts for the large majority of tokens billed.

Is it cheaper than DeepSeek V4 Flash?

Depends on the shape of your traffic. On raw input and output the GLM promo rate is lower. On cache reads, one developer measured GLM at about 4x DeepSeek V4 Flash's off-peak rate — and DeepSeek additionally halves prices off-peak. For long-context agents with high cache hit rates you need to price both separately.

Coding Plan or pay-as-you-go?

Z.ai's Coding Plan gives Flash roughly 3x the quota of GLM-5.3, which suits steady high-volume daily use. If your usage is spiky, or you also need Claude, GPT or Gemini, pay-as-you-go through one gateway is more flexible and avoids topping up separately on every platform.

Can I trust a third party that undercuts the official price?

You can use one, but check three things: whether the context ceiling is really the full 1M, whether cache billing matches, and what the rate limits are. Platforms quoting below the official promo (for example $0.06 / $0.20) are usually running a temporary subsidy that may not last.

Sources

Z.ai and bigmodel.cn official pricing documentation, the Z.ai launch announcement (2026-08-26), OpenRouter and Novita model page rates, AIHubMix and QuickSilver Pro launch posts, and a developer's cache-cost measurement posted on X (captured 2026-08-26 to 27).

Chinese models, one key for all of them

GLM-5.3, DeepSeek V4, Kimi K3 and Qwen switch on the same QCode endpoint at official rates times a service multiplier — no separate top-up on each platform.