As of 2026-09-20 · Sol now $4/$20 on promo

What does GPT-5.6 actually cost?
Tier prices, caching and real costs

The live question is the bill, not the launch sticker. Sol is now $4/$20 per the 2026-09-12 archive of the official pricing page (promotional at least through 2026-11-21; the launch sticker was $5/$30), Terra $2/$12, Luna $0.20/$1.20 per million tokens; real cost follows cache hits. Codex subscription usage is a separate meter: OpenAI’s @thsottiaux wrote on 2026-07-29 (fetched 2026-08-30) that OpenAI has not reduced usage on any subscription plans, that Sol was using Codex limits faster than expected, that usage limits were reset for all ChatGPT Work and Codex users, and that typical Sol use should last around 18% longer.

#How the bill is counted #Cache hits #Sol now $4/$20
Update 2026-07-30: Luna and Terra list prices cut

OpenAI cut Luna 80% to $0.20/$1.20 and Terra 20% to $2/$12. Sol stayed at $5/$30 as of 2026-08-02, with Fast mode (the 2026-09-12 archive of the official pricing page shows $4/$20, promotional at least through 2026-11-21). Track ongoing changes on the Luna price tracker.

Open the GPT-5.6 Luna price tracker

Official price table (per 1M tokens)

Cache writes bill at 1.25x the input rate, cache reads get a 90% discount. All three tiers share a 1,050,000-token API context and 128,000 max output; above 272K input the whole request bills at 2x input and 1.5x output. The Sol column shows current promo pricing (at least through 2026-11-21) -- the $5/$30 launch sticker appears in the examples below.

Tier Input Output Cache write (1.25x) Cache read (−90%)
gpt-5.6-sol $4 $20 $5.00 $0.40
gpt-5.6-terra $2 $12 $2.50 $0.20
gpt-5.6-luna $0.20 $1.20 $0.25 $0.02

The GA caching upgrades are worth designing for

GPT-5.6 ships more predictable prompt caching at GA: explicit cache breakpoints (you decide where the prefix splits) and a 30-minute minimum cache life. For agent workflows that means pinning system prompts and tool definitions as the cached prefix can cut long-session input costs by up to 90%.

Three reproducible cost examples

Line-by-line at official rates — swap in your own volumes to check.

Everyday agent session (Terra)

800K input + 60K output on Terra: 0.8 x $2 + 0.06 x $12 = $1.60 + $0.72 = about $2.32 (short-context list rates as of 2026-08-02). If 80% of the input is a cached read ($0.20/M), the input part is 0.16 x $2 + 0.64 x $0.20 = $0.448, so about $1.17 total.

Hard task (Sol, high cache hit)

500K input (400K of it cached reads) + 50K output at Sol's current price per the 2026-09-12 archive of the official pricing page ($4 input, $0.40 cached read, $20 output; promotional at least through 2026-11-21): 0.1 x $4 + 0.4 x $0.40 + 0.05 x $20 = $0.40 + $0.16 + $1.00 = about $1.56 (about $2.20 at the $5/$30 launch price). Caching brings a Sol deep-dive close to Terra's uncached price for the same volume (about $1.60).

Bulk light tasks (Luna)

1,000 calls at 10K input + 1K output each (10M input + 1M output): 10 x $0.20 + 1 x $1.20 = about $3.20 on Luna after the 2026-07-30 cut. The same volume on Sol runs about $60 at the current $4/$20 (about $80 at the $5/$30 launch price) — moving light tasks down to Luna is still the most direct saving.

How the prices compare

Sol ($5/$30, 2026-08 basis) matches Claude Opus 4.8 ($5/$25) on input and runs slightly higher on output; Terra ($2/$12 after the 2026-07-30 cut) is the everyday default; Luna ($0.20/$1.20) is about 1/25 of Sol list price for high-volume work. As of 2026-08-02, OpenAI list rates separate flagship capability from flagship spend.

Prices and usage on QCode

QCode's /models page shows live rates matching OpenAI's official prices, and usage inside your plan is metered per token at those rates; a request with more than 272K input tokens bills the whole request at the long-context rate, matching OpenAI's rule. The three tiers share one API key and one quota with GPT-5.5, Claude and Chinese models — low-latency access from mainland China with local payment, on plans from $8.57/mo.

Pricing FAQ

How exactly is caching billed?

Cache writes cost 1.25x the input rate (Sol $6.25, Terra $2.50, Luna $0.25 per 1M as of 2026-08-02). Cache reads stay about 10% of input (Sol $0.50, Terra $0.20, Luna $0.02).

Is long context billed extra?

There is a long-context tier. OpenAI's Sol model page says that once input exceeds 272K, the whole request bills at 2x input and 1.5x output ("Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request."). Note 272K is not a ceiling on what you can send -- the window stays 1,050,000 -- it is a price cliff. The lever is still cache hits on a stable prefix, not truncating context. QCode bills by this same rule.

Does switching tiers affect billing or quota?

No. The tiers are just different model ids, each billed at its own rate against the same quota. You can mix them per request in one project — say Sol to plan, Terra to execute, Luna for batch work.

Do QCode prices match OpenAI's official rates?

Sol and Terra, yes. The /models page shows those two tiers at official rates (Sol $5/$30, Terra $2/$12 as of 2026-08-02); QCode does not currently support gpt-6-luna or gpt-5.6-luna (both removed from /models); please use another GPT model, such as gpt-6.1-sol, gpt-6-sol or gpt-5.6-terra.

Use GPT-5.6 at official rates

GPT-5.6 Sol and Terra are included in a QCode monthly plan, billed per token with per-model prices on /models — from $8.57/mo, month to month, cancel whenever. QCode does not currently support gpt-6-luna or gpt-5.6-luna (both removed from /models); please use another GPT model, such as gpt-6.1-sol, gpt-6-sol or gpt-5.6-terra.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.