Pricing tracker · as of 2026-10-07

1M Long-Context Pricing Compared (2026): Who Charges More, and From Where

As of 2026-10-07, Anthropic's pricing page states that its 1M context is billed at standard pricing: on Claude Opus 5.5 and Sonnet 5.5 a 900K-token request costs the same per token as a 9K-token one. OpenAI prices the whole request at 2x input and cache rates and 1.5x output once GPT-6 Sol or GPT-6.1 Sol receives more than 272K input tokens; xAI moves every token in a Grok 4.7 request to its higher tier once the prompt reaches 200K; DeepSeek's official price list for V4.1 Flash shows no long-context surcharge tier (as of 2026-10-07), only peak and off-peak rates. Three worked examples converted from the official price lists follow.

Updated 2026-10-07

#1M context#272K threshold#200K threshold#Long-context surcharge

Four numbers that matter

Standard rate

Claude at 1M context

Per the official pricing page, Claude 4.6 and later models get the full 1M context at standard pricing, so a 900K-token request is billed at the same per-token rate as a 9K-token one. Opus 5.5 is $4 input / $20 output, Sonnet 5.5 is $2 / $10.

>272K

GPT-6 Sol and GPT-6.1 Sol threshold

Above 272K input tokens the full request is priced at 2x input and cache rates and 1.5x output: GPT-6.1 Sol goes from $2 / $10 to $4 / $15.

≥200K

Grok 4.7 threshold

Once the prompt reaches 200K, every token in the request is billed at the higher tier: $2 / $6 becomes $4 / $12. Grok 4.7's context is 500K, not 1M.

$0.3 / $1.2

DeepSeek V4.1 Flash peak rate

1M context; cache-miss input and output cost $0.3 / $1.2 per million tokens at peak and half that, $0.15 / $0.6, off-peak. The official price list shows no long-context surcharge tier (as of 2026-10-07).

How the four vendors' price lists handle long context

As of 2026-10-07, the four vendors' official pages fall into three groups on long-context pricing. Stated as no surcharge: Anthropic's pricing page says Claude 4.6 and later models include the full 1M token context window at standard pricing, and its context-window docs list Opus 5.5 and Sonnet 5.5 as 1M models where 1M is the default, no beta header is needed and long-context requests are billed at standard pricing. A stated threshold that reprices the whole request: OpenAI's GPT-6 Sol and GPT-6.1 Sol switch above 272K input tokens, xAI's Grok 4.7 once the prompt reaches 200K, and in both cases every token in the request moves to the higher tier, not just the part over the line. No long-context surcharge tier listed in the official price list: DeepSeek V4.1 Flash has a 1M context, its price list shows only peak and off-peak rates, and the official formula is number of tokens × price.

Price-list changes since September

On 2026-09-29 the OpenAI changelog announced GPT-6.1 Sol, stating that standard pricing for prompts with up to 272K input tokens is $2 input, $0.10 cached input, $2.50 cache write and $10 output; the full-request rule above 272K on its model page is word for word the one on GPT-6 Sol, released 2026-09-22. Anthropic released Opus 5.5 on 2026-09-22 and Sonnet 5.5 on 2026-09-28, both with a 1M context and up to 128K output tokens per request. xAI's release notes give Grok 4.7 a 500K context window, priced at $2 / $0.50 / $6 below 200K prompt tokens and $4 / $1 / $12 above. DeepSeek released V4.1 Flash on 2026-09-10 and lowered API prices accordingly; its peak and off-peak tiers have applied since 16:00 UTC on 2026-08-16.

Timeline

2026-09-10

DeepSeek releases V4.1 Flash with a 1M context; API prices drop accordingly, still split into peak and off-peak tiers.

2026-09-22

Claude Opus 5.5 and GPT-6 Sol launch on the same day: the first bills its 1M context at standard pricing, the second reprices the whole request above 272K input tokens.

2026-09-29

GPT-6.1 Sol launches: $2 input / $10 output up to 272K, $4 / $15 above it, with cached input rising from $0.10 to $0.20.

Confirmed vs not spelled out

Confirmed (verbatim on the official pages)

Everything below can be checked word for word on the official pages: Claude 4.6 and later models get the full 1M context at standard pricing, a 900K request is billed at the same per-token rate as a 9K one, and caching and Batch discounts apply across the whole window; Opus 5.5 is $4 / $20 and Sonnet 5.5 $2 / $10, with Batch at $2 / $10 and $1 / $5; GPT-6 Sol and GPT-6.1 Sol price the full request at 2x input and cache rates and 1.5x output above 272K input tokens, and both have a 1,050,000-token context window with a 922,000-token maximum input; Grok 4.7 has a 500K context and two price rows around 200K; DeepSeek V4.1 Flash has a 1M context, peak and off-peak rates and published peak hours. The per-request totals in the examples are our own conversion from the official price lists, not figures published by any vendor.

Not spelled out by the vendors

Three points are not settled by the official pages, and this page does not settle them on the vendors' behalf: first, OpenAI does not say whether 272K means exactly 272,000 tokens, nor whether cached tokens count toward the threshold, so the examples here use 270,000 and 280,000 to stay clear of the boundary; second, xAI's table says ≥ 200k and its footnote says reaches, while the release notes say above and the Fast section says exceeds, so for a request of exactly 200K tokens go by what you are actually charged; third, DeepSeek's official price list shows no long-context surcharge tier (as of 2026-10-07), so the examples use its listed rates, and the same page states that DeepSeek reserves the right to adjust prices.

How to choose: start with where your inputs usually land

Inputs mostly between 200K and 272K

In this band Grok 4.7 is already on its higher tier, GPT-6 Sol and GPT-6.1 Sol are still on standard rates, Claude is stated to stay at standard pricing, and DeepSeek's official price list shows no long-context surcharge tier. Converted from the official price lists, one request of 250,000 input + 8,000 output tokens costs $0.58 on both Sonnet 5.5 and GPT-6.1 Sol, $1.16 on Opus 5.5, $1.096 on Grok 4.7 and $0.0846 on DeepSeek V4.1 Flash at peak.

Inputs often above 272K

Past 272K, OpenAI moves to its higher tier as well; Claude is stated to stay at standard pricing, and DeepSeek's official price list shows no long-context surcharge tier, so its listed rates are used. Converted from the official price lists, one request of 400,000 input + 8,000 output tokens costs $0.88 on Sonnet 5.5, $1.76 on Opus 5.5, $1.72 on GPT-6.1 Sol, $1.696 on Grok 4.7, and $0.1296 at peak or $0.0648 off-peak on DeepSeek V4.1 Flash. At that point GPT-6.1 Sol's input rate ($4) matches Opus 5.5's.

Three worked examples you can recompute

All figures are converted from the official price lists. First, 400,000 input + 8,000 output tokens with no cache hits: Sonnet 5.5 = 0.4 × $2 + 0.008 × $10 = $0.88; Opus 5.5 = 0.4 × $4 + 0.008 × $20 = $1.76; GPT-6.1 Sol and GPT-6 Sol on the long-context tier = 0.4 × $4 + 0.008 × $15 = $1.72; Grok 4.7 on the 200K-plus tier = 0.4 × $4 + 0.008 × $12 = $1.696; DeepSeek V4.1 Flash at peak = 0.4 × $0.3 + 0.008 × $1.2 = $0.1296, or $0.0648 off-peak. Second, GPT-6.1 Sol's 272K threshold: 270,000 input + 8,000 output = 0.27 × $2 + 0.008 × $10 = $0.62, while 280,000 input + 8,000 output = 0.28 × $4 + 0.008 × $15 = $1.24, so 10,000 more input tokens double the bill; the same two requests on Sonnet 5.5 cost $0.62 and $0.64. Third, Grok 4.7's 200K threshold: 250,000 input + 8,000 output = 0.25 × $4 + 0.008 × $12 = $1.096, against $0.548 at the below-200K rates. None of this includes Batch, caching or regional premiums.

On QCode

QCode bills per token; see /models for each model's price. Of the models on this page, claude-opus-5-5, claude-sonnet-5-5, gpt-6.1-sol, gpt-6-sol and deepseek-v4.1-flash can be called with the same key; Grok 4.7 is not among them. The examples here are conversions from each vendor's official price list, not a QCode quote and not a description of how QCode bills.

Frequently asked questions

Do Claude Opus 5.5 and Sonnet 5.5 charge extra for the full 1M context?

No. The official pricing page says Claude 4.6 and later models include the full 1M token context window at standard pricing, and gives the example that a 900k-token request is billed at the same per-token rate as a 9k-token request; the context-window docs add that 1M is the default, you don't need a beta header, and long-context requests are billed at standard pricing. Caching and Batch discounts apply across the full window, and at any request length the standard rates stay at $4 / $20 for Opus 5.5 and $2 / $10 for Sonnet 5.5.

Does GPT-6.1 Sol only charge more for the tokens above 272K?

No, the whole request is repriced. The model page reads: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request. Per the official price list, GPT-6.1 Sol costs $2 input, $0.10 cached input, $2.50 cache writes and $10 output up to 272K, and $4, $0.20, $5 and $15 above it; GPT-6 Sol differs only in cached input ($0.20 and $0.40). Both have a 1,050,000-token context window and a 922,000-token maximum input.

Does Grok 4.7 have a 1M context, and how is it billed above 200K?

No, xAI's price table lists Grok 4.7 with a 500K context. Pricing comes in two rows: below 200K prompt tokens it is $2 input, $0.50 cached input and $6 output; at 200K and above it is $4, $1 and $12, and the footnote says the higher rate then applies to all tokens in the request. The xAI models listed with a 1M context are grok-4.3 and the grok-4.20 family, whose 200K-plus rows are likewise double the lower rows (grok-4.3 goes from $1.25 / $2.50 to $2.50 / $5.00).

How is DeepSeek V4.1 Flash billed across its 1M context?

Its official price list shows no long-context surcharge tier (as of 2026-10-07), only peak and off-peak rates. V4.1 Flash (API model name deepseek-flash) has a 1M context, and per million tokens it costs $0.006 at peak and $0.003 off-peak for cache-hit input, $0.3 and $0.15 for cache-miss input, and $1.2 and $0.6 for output. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays; all other hours, including weekends and Chinese public holidays in full, are off-peak.

Do Batch and regional endpoints stack with long-context pricing?

Each vendor words it differently. Anthropic says caching and Batch discounts apply at standard rates across the full context window, with Batch at $2 / $10 for Opus 5.5 and $1 / $5 for Sonnet 5.5, and US-only inference via inference_geo adds a 1.1x multiplier to all token categories. OpenAI's Batch table puts GPT-6.1 Sol's above-272K tier at $2 input and $7.50 output, and regional processing adds a 10% premium where offered. xAI's Batch discount list does not include Grok 4.7, and models not on it get no Batch discount; its US regional endpoint bills 1.1x including long-context rates, which is $4.40 / $1.10 / $13.20 for Grok 4.7 above 200K. DeepSeek's official price list shows no Batch or regional rates (as of 2026-10-07).

Can I call these models on QCode?

claude-opus-5-5, claude-sonnet-5-5, gpt-6.1-sol, gpt-6-sol and deepseek-v4.1-flash can be called with the same key; Grok 4.7 cannot. QCode bills per token, and each model's price is on /models. Every figure on this page is converted from the vendors' official price lists and is not a QCode quote.

Sources

Anthropic: the official pricing page (model price table, long-context pricing, Batch and data-residency rates), the context-window docs, the models overview, and the Opus 5.5 and Sonnet 5.5 model pages (release dates). OpenAI: the developer docs pricing page (Standard and Batch tables, short- and long-context definitions), the GPT-6.1 Sol and GPT-6 Sol model pages (the 272K rule, context window and maximum input) and the API changelog (release dates and the up-to-272K standard rates). xAI: the docs Models page, the Pricing page (Batch, US regional endpoint and Fast sections) and the Release Notes. DeepSeek: the API docs Models & Pricing page and the Change Log. All retrieved on 2026-10-07.

One key: switch the model and compare

QCode bills per token, with each model's price on /models; claude-sonnet-5-5, gpt-6.1-sol and deepseek-v4.1-flash all run on the same key.

Related reading

The figures and quotes on this page were checked on 2026-10-07; the vendors' official pages take precedence, and prices may change without notice. Every per-request total is a conversion from the official price lists, not a quote or a promise. Model availability is whatever /models shows.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.