Prompt Caching Prices in 2026: How Claude, GPT and DeepSeek Charge and How to Save
As of 2026-10-07, cache reads on Claude Opus 5.5 and Sonnet 5.5 both cost $0.20 per million tokens; a 5-minute cache write is 1.25x the base input price ($5 / $2.50) and a 1-hour write is 2x ($8 / $4). GPT-6.1 Sol charges $0.10 for cached input and GPT-6 Sol $0.20, with cache writes at $2.50 on both. DeepSeek-V4-Pro cache hits cost $0.022 off-peak and $0.044 at peak. This page puts the three vendors' billing rules, TTLs and minimum cache lengths side by side, with two worked examples converted from the official price lists.
Updated 2026-10-07
Four numbers to remember
Opus 5.5 / Sonnet 5.5 cache reads (per million tokens)
Both are $0.20. The price-list footnote bills Opus 5.5 cache hits and refreshes at 0.05x the base input price (5% of $4); Sonnet 5.5 uses the standard 0.1x (10% of $2).
Claude cache writes: 5-minute / 1-hour
From the official multiplier table: a 5-minute cache write is 1.25x the base input price, a 1-hour write 2x. In unit prices that is $5 / $8 on Opus 5.5 and $2.50 / $4 on Sonnet 5.5.
GPT-6.1 Sol / GPT-6 Sol cached input
Standard rates for prompts of up to 272K input tokens. Both have $2 input and $2.50 cache writes (1.25x); cached input is 5% of the uncached rate on GPT-6.1 Sol and 10% on GPT-6 Sol.
DeepSeek-V4-Pro cache hits: off-peak / peak
The official price list splits input into a cache-hit and a cache-miss rate; a miss costs $0.66 / $1.32. DeepSeek-V4.1-Flash cache hits cost $0.003 / $0.006.
How caching is billed
As of 2026-10-07, prompt caching works the same basic way at all three vendors: when a repeated prefix hits the cache, those tokens are billed at a discounted rate. What differs is whether writes carry a surcharge, how large the read discount is and how long the cache lives. Anthropic: the default cache lifetime is 5 minutes and every hit refreshes it at no extra cost; you can choose 1 hour instead, which raises the write price from 1.25x to 2x the base input price. In the vendor's own arithmetic, the 5-minute cache pays off after one read and the 1-hour cache after two. OpenAI: the official docs say prompt caching is enabled by default for supported models; on GPT-5.6 and later, cache writes cost 1.25x the uncached input rate, and reads cost 0.1x on most models and 0.05x on GPT-6.1 Sol. The write price is not an additive fee: each input token is billed at exactly one of the uncached, cached or cache-write rates. DeepSeek: the vendor says its disk cache is on by default for all users with no code changes, and the price list simply has a cache-hit and a cache-miss rate.
The vendors' framing: cache reads at just 0.05x on Opus 5.5 and GPT-6.1 Sol
Right now both Anthropic and OpenAI have models with a cache-read multiplier below the usual 0.1x. Anthropic's price-list footnote bills Opus 5.5 cache hits and refreshes at 0.05x the base input price — $4 input, $0.20 cache reads; most other models use the standard 0.1x, and Fable 5.1 uses 0.025x. OpenAI's caching guide says that on GPT-5.6 and later, cache writes cost 1.25x the uncached input rate and reads cost 0.1x on most models and 0.05x on GPT-6.1 Sol; for earlier models, the official comparison table lists no additional cache-write charge. DeepSeek's price list, by contrast, has just two input rates: cache hit and cache miss.
Three dates that matter for cache pricing
Claude Opus 5.5 launches at $4 / $20 (cache reads $0.20 on the current price list); the same day OpenAI releases GPT-6 Sol at $2 input, $0.20 cached input and $10 output for prompts of up to 272K input tokens.
Claude Sonnet 5.5 launches; its row on the current price list reads input $2, 5-minute cache write $2.50, 1-hour cache write $4, cache reads $0.20, output $10.
OpenAI releases GPT-6.1 Sol: for prompts of up to 272K input tokens it costs $2 input, $0.10 cached input, $2.50 cache write and $10 output, with cached input at 5% of the uncached rate.
Confirmed vs not verified
Confirmed (verbatim in the official pages)
All of the following can be checked word for word in the official docs and price lists: cache reads of $0.20 on Opus 5.5 and Sonnet 5.5, 5-minute writes of $5 and $2.50, 1-hour writes of $8 and $4; the 0.05x cache-read multiplier on Opus 5.5; Claude's default 5-minute cache lifetime, counted from the start of the request; a 512-token minimum cacheable length on Opus 5.5 and Sonnet 5.5 and 1,024 tokens on Sonnet 5, with shorter prompts processed without caching and no error returned; cached input of $0.10 on GPT-6.1 Sol and $0.20 on GPT-6 Sol, with cache writes of $2.50 on both; a 1,024-token minimum and a cache lifetime of at least 30 minutes on GPT-5.6 and later; above 272K input tokens, GPT-6 Sol and GPT-6.1 Sol bill input and cache at 2x and output at 1.5x for the full request; the DeepSeek-V4-Pro and DeepSeek-V4.1-Flash hit and miss rates and the peak hours. The two worked examples on this page are conversions from those official prices, not measured bills.
Unverified and not published by the vendor
Three things have not been verified, so don't budget around them: first, DeepSeek publishes no fixed cache lifetime — the vendor only says the cache works on a best-effort basis with no guaranteed hit rate and is usually cleared within a few hours to a few days once it is no longer used; second, OpenAI says cached input is discounted up to 95%, and the example deployments in its docs report hit rates of about 70% and above 90% — those are the vendor's illustrations, not a promise about your workload; third, the cache-hit-rate screenshots and comparison figures circulating online have no first-hand source, and this page does not carry them.
Two worked examples you can recompute
Example 1: a long coding-agent session, a 100K prefix repeated over 20 turns
Converted from the official price lists, with fixed inputs: a 100K-token fixed prefix (system prompt, tool definitions, project context), 20 consecutive turns each less than 5 minutes apart, counting only that 100K prefix and not each turn's new content or output. Sonnet 5.5: uncached 20 × 0.1M × $2 = $4.00; cached, one 5-minute write at $0.25 plus 19 reads at 19 × $0.02 = $0.38, $0.63 in total. Opus 5.5: $8.00 vs $0.88 (write $0.50 + reads $0.38). GPT-6 Sol: $4.00 vs $0.63; GPT-6.1 Sol: $4.00 vs $0.44 (write $0.25 + reads 19 × $0.01). When comparing across vendors, note that these figures assume equal token counts, while the same text may tokenize differently on each vendor's models; Anthropic's price list, for example, says the newer tokenizer in Claude 4.7 and later produces about 30% more tokens for the same text than the previous one, with the exact increase depending on the content.
Example 2: a 5-minute or a 1-hour TTL
Converted from the official price lists, with fixed inputs: Sonnet 5.5, a 100K prefix, 6 turns, 10 minutes between turns. Uncached: 6 × $0.20 = $1.20. With the 5-minute TTL the cache has expired every time and is written again, 6 × $0.25 = $1.50, which is more than not caching at all. With the 1-hour TTL, one write at $0.40 plus 5 reads at $0.10 comes to $0.50. Under the same conditions Opus 5.5 costs $2.40, $3.00 and $0.90. The vendor's advice: if calls always come less than 5 minutes apart, stay on the 5-minute cache, since every hit renews it at no extra charge; move to the 1-hour cache when the gaps fall between 5 minutes and an hour. For comparison, on OpenAI GPT-5.6 and later a cached prefix stays eligible for reuse for at least 30 minutes after its latest write or reuse, so a 10-minute gap is still inside that window under the official rules.
How to turn it on and save: five steps
First, on the Claude API, a single cache_control field at the top level of the request turns on automatic caching: the system places the breakpoint on the last cacheable block and moves it forward as the conversation grows. For one hour, set ttl to 1h inside cache_control; a prompt can carry up to 4 breakpoints. Second, mind the order: the cached prefix runs tools, then system, then messages, and a change at any level invalidates that level and everything after it, so put stable tool definitions, system prompts and long documents first and timestamps or other changing content last. Third, in Claude Code with an API key, the main conversation defaults to a five-minute TTL; the official docs say to set promptCacheTtl to 1h (or use the CLAUDE_CODE_PROMPT_CACHE_TTL environment variable), which needs Claude Code v2.1.242 or later. After a switch with /model, the first request re-reads the whole conversation uncached. Fourth, on OpenAI GPT-5.6 and later, implicit mode, which the vendor says works out of the box for most use cases, places a breakpoint at the end of the latest eligible message; in explicit mode, content after the last breakpoint is billed at the uncached rate with no cache-write charge, so you can avoid writing content that changes every time and is unlikely to be reused. Fifth, check the usage fields: on Claude, cache_creation_input_tokens and cache_read_input_tokens (if both are 0, nothing was cached); on OpenAI, usage.input_tokens_details.cached_tokens and cache_write_tokens; on DeepSeek, prompt_cache_hit_tokens and prompt_cache_miss_tokens.
On QCode
QCode bills per token; each model's unit prices are listed at /models. The figures on this page come from the vendors' official price lists and are there to explain how caching is priced and how to save; for what is actually charged on QCode, go by /models and the per-call model details in the console.
Frequently asked questions
How much do Claude cache reads cost?
As of 2026-10-07, cache reads cost $0.20 per million tokens on both Claude Opus 5.5 and Sonnet 5.5. The multipliers differ: Opus 5.5 is billed at 0.05x the base input price (5% of $4), Sonnet 5.5 at the standard 0.1x (10% of $2), and Fable 5.1 at 0.025x ($0.25). The price-list column is called Cache hits and refreshes: reading a prompt prefix from the cache, which also refreshes its lifetime.
What does a 1-hour Claude cache write cost, and is it worth it?
A 1-hour cache write is 2x the base input price: $4 per million tokens on Sonnet 5.5 and $8 on Opus 5.5 (the 5-minute tier is $2.50 and $5). The vendor's arithmetic at the 0.1x read rate: the 5-minute cache pays off after one read, the 1-hour cache after two. Converted from the official price list, a 100K prefix on Sonnet 5.5 used for 6 turns 10 minutes apart costs $0.50 on the 1-hour tier and $1.50 on the 5-minute tier, which expires and is written again every turn. If calls always come less than 5 minutes apart, the vendor recommends staying on the 5-minute tier.
Do I have to turn on OpenAI caching, and are writes charged?
There is nothing to turn on: the official docs say prompt caching is enabled by default for supported models. GPT-5.6 and later charge for writes at 1.25x the uncached input rate ($2.50 on both GPT-6 Sol and GPT-6.1 Sol), but it is not an additive fee: each input token is billed at only one of the uncached, cached or cache-write rates. The lifetime is controlled with prompt_cache_options.ttl, whose only supported value, 30m, is also the default; reuse refreshes the lifetime without another write charge. The minimum cacheable length is 1,024 tokens.
How much do DeepSeek cache hits cost?
As of 2026-10-07, per million input tokens on the official price list: DeepSeek-V4-Pro cache hits cost $0.022 off-peak and $0.044 at peak, misses $0.66 / $1.32; DeepSeek-V4.1-Flash (API model name deepseek-flash) hits cost $0.003 / $0.006, misses $0.15 / $0.3. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays, and off-peak rates are half the peak rates. Converted from the official price list, a 100K prefix over 20 turns on V4-Pro at peak costs $2.64 with no hits, or $0.2156 with a miss on the first turn and hits on the next 19.
Do cache hits count against rate limits?
The two vendors differ. Among the use cases for its 1-hour cache, Anthropic states that cache hits are not deducted against your rate limit. OpenAI's FAQ says cached input tokens still count toward tokens-per-minute limits and that prompt caching does not change how rate limits are calculated.
How is caching billed on QCode?
QCode bills per token; each model's unit prices are listed at /models. The figures on this page come from the vendors' official price lists. You can check the model, token count and cost of every call in the model call details on the usage statistics page of the console.
Sources
Claude pricing and caching rules: Anthropic's official pricing page and Prompt caching docs (crawled 2026-10-07, covering per-model cache read and write prices, multipliers and footnotes, TTLs, minimum cacheable lengths, invalidation rules and usage fields); release dates: Anthropic's official release notes. Claude Code TTL settings: the prompt caching and cost management pages in the official Claude Code docs. OpenAI: the official pricing page, the GPT-6 Sol and GPT-6.1 Sol model pages, the Prompt caching guide and the API changelog (crawled the same day). DeepSeek: the official Models & Pricing page and the Context Caching guide (crawled the same day). QCode: the billing and cost optimization pages on docs.qcode.cc.
Billed per token, prices on /models
QCode bills per token; each model's unit prices are listed at /models. For how caching is priced and how to save, start with the official prices and the two worked examples on this page.
Related reading
Save on AI coding costs with Claude Code
Use Claude Code the smart way: master token optimization and choose the most cost-effective plan.
Claude Opus 5.5 complete guide
Pricing, caching, the comparison with Opus 5 and calling it on QCode.
GPT-6 Sol / Luna pricing tracker
Official rates, billing rules and recomputable examples for two OpenAI reasoning models.
The prices and rules on this page were checked on 2026-10-07; the official Anthropic, OpenAI and DeepSeek pages take precedence, and upstream changes may happen without notice. The two worked examples are illustrations converted from the official price lists: they count only the input cost of the repeated prefix, exclude new content and output, and are not measured bills. Discount percentages and example hit rates quoted from the vendors are their own claims. Model availability is whatever /models shows.