GPT-6.1 Sol API Guide: Pricing, Specs and GPT-6 Sol Compared
As of 2026-10-07, GPT-6.1 Sol (model ID gpt-6.1-sol) is the new Sol model OpenAI released at DevDay on 2026-09-29, with standard prices of $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million tokens, a 1,050,000-token context window, 128,000 max output tokens and an April 30, 2026 knowledge cutoff. Against GPT-6 Sol, only one of the four list prices changes: cached input is halved ($0.20 to $0.10). Requests with more than 272K input tokens are billed at long-context rates for the whole request. This page compares the two line by line from the official model pages and price list, and covers when to pick it.
Updated 2026-10-07
Four key numbers
Input / output (per million tokens)
Official standard rates for prompts up to 272K input tokens, the same as gpt-6-sol. OpenAI's launch post calls this one-fifth of GPT-6 Astra's standard input and output prices (Astra is $10 / $50).
Cached input / cache writes
The official model page prices cached input at 5% of the uncached input rate and cache writes at 1.25x. On gpt-6-sol cached input is $0.20 (10%); cache writes are the same $2.50.
Context / max output
From the official model page: a 1,050,000-token context window, 922,000 max input tokens and 128,000 max output tokens, with an April 30, 2026 knowledge cutoff. All three limits match gpt-6-sol.
Long-context threshold
A prompt with more than 272K input tokens is billed at 2x input and cache rates and 1.5x output for the full request, i.e. $4 / $0.20 / $5 / $15, not just on the tokens above the threshold.
What GPT-6.1 Sol is
As of 2026-10-07, GPT-6.1 Sol is the tier OpenAI's model catalog recommends to balance intelligence and cost; its model ID and default snapshot are both gpt-6.1-sol. OpenAI's launch post calls it an upgrade to GPT-6 Sol that nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's standard input and output prices. It takes text and image input and produces text, with no audio or video. The Responses, Chat Completions and Batch endpoints are supported, but tool calling requires the Responses API; Chat Completions accepts requests without tools only. reasoning.effort accepts low, medium (default), high, xhigh and max; none and minimal are not supported.
The vendor's own framing (DevDay, 2026-09-29)
The OpenAI API changelog entry for 2026-09-29 reads: GPT-6.1 Sol (gpt-6.1-sol) released for complex coding and professional work at a lower cost than GPT-6 Astra; standard pricing for prompts up to 272K input tokens is $2 input, $0.10 cached input, $2.50 cache write and $10 output, and it supports Multi-agent in beta in the Responses API. The DevDay 2026 page from the same day lists it under API and says to use it in the API and Codex, where available. The launch post says cached input at $0.10 is "95% less than standard input pricing and 50% less than GPT-6 Sol's cached input pricing". The post also carries OpenAI's own benchmark scores, which this page does not transcribe; it adds that GPT-6 Astra should still be used for the most difficult scientific research tasks.
The timeline
GPT-6 Sol (gpt-6-sol) is released at standard rates of $2 input, $0.20 cached input and $10 output.
At DevDay 2026, GPT-6.1 Sol (gpt-6.1-sol) is released: cached input drops to $0.10, while input, cache writes and output stay at GPT-6 Sol's prices.
gpt-6.1-sol becomes callable on QCode. As of 2026-10-07 both gpt-6.1-sol and gpt-6-sol run on the same key.
Confirmed vs not published
Confirmed (verbatim in the official pages)
Each of the following can be checked word for word on OpenAI's official model pages, price list, changelog and launch post: released 2026-09-29; model ID gpt-6.1-sol; 1,050,000 context, 922,000 max input, 128,000 max output; knowledge cutoff April 30, 2026; standard input $2, cached input $0.10, cache writes $2.50, output $10; above 272K input tokens the full request is billed at 2x input and cache rates and 1.5x output; Batch and Flex 50% below Standard, Fast at 2x Standard; regional processing adds 10% where available; the none and minimal reasoning efforts are not supported; tool calling requires the Responses API. The "near-Astra" claim and the scores behind it are OpenAI's own statements and testing. On the QCode side: gpt-6.1-sol has been callable on QCode since 2026-09-30.
Unverified and not published by the vendor
As of 2026-10-07, these points have no official answer, so don't plan around them: first, GPT-6.1 Sol Ultrafast: the launch post says "in the coming days", the Codex docs say "coming later", and the Ultrafast table on the official price list still lists only gpt-6-astra, with no 6.1 Sol price; second, retirement dates: the official deprecations page lists neither gpt-6.1-sol nor gpt-6-sol for retirement, and gpt-6-sol is in fact the recommended replacement for gpt-5.1 and gpt-5.3-codex in the 2026-10-01 notice; third, the "near-Astra" claim and all scores come only from OpenAI's own testing, and this page found no third-party replication; fourth, comparison screenshots and multiplier math circulating online are left out unless a first-hand source carries them.
How it compares with GPT-6 Sol and GPT-6 Astra
vs GPT-6 Sol (gpt-6-sol)
Line by line from the official price list (Standard, up to 272K input): input is $2 on both, cache writes $2.50 on both, output $10 on both; only cached input differs, $0.10 on gpt-6.1-sol vs $0.20 on gpt-6-sol. Above 272K it is $0.20 vs $0.40, and on Batch $0.05 vs $0.10. Context 1,050,000, max input 922,000 and max output 128,000 are identical; the knowledge cutoff is April 30 vs April 20, 2026. The behavioral difference is in parameters: gpt-6-sol supports the none reasoning effort and allows function calling in Chat Completions only with reasoning_effort set to none; gpt-6.1-sol does not support none, and tool calling always goes through Responses.
vs GPT-6 Astra (gpt-6-astra)
The Astra row of the official price list is $10 input, $1.00 cached input, $12.50 cache writes and $50 output. Converted from the official price list, gpt-6.1-sol is exactly one-fifth of Astra on input and output and one-tenth on cached input. OpenAI's model-selection guide suggests considering GPT-6.1 Sol for complex projects where cost matters and comparing it with Astra on the same task to weigh quality against cost; OpenAI also says Astra should still be used for the most difficult scientific research tasks.
Three things to check when moving from gpt-6-sol
OpenAI's GPT-6 guide says that if you already use gpt-6-sol, you should review the migration guidance before switching to GPT-6.1 Sol. Per the official checklist: first, reasoning effort: GPT-6.1 Sol does not support none, so use low instead; if you used minimal, start with low and compare results on representative tasks. Second, tool calling: GPT-6.1 Sol supports Chat Completions, but tool calling requires the Responses API. Third, parameters: when reasoning effort is not none, remove temperature, top_p and top_logprobs. Also budget for long jobs that may cross 272K input tokens: converted from the official price list, 270K input plus 10K output (no caching) costs $0.64, while 280K input plus 10K output costs $1.27, so 10K more input nearly doubles the whole request.
Calling it on QCode
gpt-6.1-sol has been callable on QCode since 2026-09-30, and gpt-6-sol is callable too; one API key covers both. QCode bills per token; per-model prices are on /models. Paths follow the QCode docs: for the OpenAI Python/JS SDK, set base_url to https://api.qcode.cc/openai/v1, which the SDK turns into /openai/v1/chat/completions; the OpenAI Responses path is /openai/v1/responses, and Codex CLI must use it: in config.toml set base_url to https://api.qcode.cc/openai and wire_api to responses, and the model field can take gpt-6.1-sol from the docs' model list. Users in North America and Europe can also use us.qcode.cc (Los Angeles).
Frequently asked questions
How much does the gpt-6.1-sol API cost?
As of 2026-10-07, the official standard rates (up to 272K input tokens) per million tokens are $2 input, $0.10 cached input, $2.50 cache writes and $10 output. Above 272K input tokens the full request moves to $4 / $0.20 / $5 / $15. Batch and Flex are $1 / $0.05 / $1.25 / $5, Fast is $4 / $0.20 / $5 / $20, and regional processing adds 10% where available.
What is the difference between gpt-6.1-sol and gpt-6-sol?
On list price, only cached input differs: $0.10 vs $0.20 (5% vs 10% of the uncached input rate); input, cache writes and output are the same. Context, max input and max output are identical, and the knowledge cutoff is ten days later (April 30 vs April 20, 2026). On parameters, gpt-6.1-sol does not support the none reasoning effort and tool calling goes through Responses. OpenAI says it delivers substantial improvements over GPT-6 Sol across complex professional tasks; that is the vendor's own claim.
How are prompts over 272K input tokens priced?
For the full request, not just the excess. The official model page says: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request. The price list defines short context as up to 272K input tokens and long context as more than 272K. Converted from the official price list: 270K input plus 10K output (no caching) is $0.64; 280K input plus 10K output is $1.27. The same rule applies to gpt-6-sol.
What are the context window, max output and knowledge cutoff?
From the official model page: a 1,050,000-token context window, 922,000 max input tokens, 128,000 max output tokens and an April 30, 2026 knowledge cutoff. Input can be text or images and output is text; audio and video are not supported.
When should I choose gpt-6.1-sol?
OpenAI suggests considering GPT-6.1 Sol for complex projects where cost matters and comparing it with Astra on the same task; the Codex docs also recommend GPT-6.1 Sol for complex coding and agentic workflows when available. If you already run gpt-6-sol, the only list-price difference is cached input, and OpenAI says the lower cache price gives more room to run agents that reuse context across requests; converted from the official price list, you pay $0.10 less per million cached input tokens. Requests that rely on the none reasoning effort, including function calling in Chat Completions with none, need reworking before you switch.
Can I call gpt-6.1-sol on QCode right now?
Yes. gpt-6.1-sol has been callable on QCode since 2026-09-30, and gpt-6-sol is callable too, both on the same key; when moving from gpt-6-sol, first run through the three migration checks above (reasoning effort, tool calling, parameters). QCode bills per token, and per-model prices and the current model list are on /models.
Sources
Specs and unit prices: the gpt-6.1-sol and gpt-6-sol model pages in OpenAI's developer docs and the official pricing page (the Standard, Batch, Flex, Fast and Ultrafast tables and the 272K threshold note). Release date and contents: the OpenAI API changelog (the 2026-09-29 and 2026-09-22 entries) and the DevDay 2026 page on learn.chatgpt.com. Positioning and vendor claims: OpenAI's launch post Introducing GPT-6.1 Sol (read from a Wayback Machine capture). Migration and model choice: the official GPT-6 guide, the model-selection guide, the Codex models docs and the deprecations page. All crawled 2026-10-07. The QCode setup comes from two docs.qcode.cc pages, Endpoints and API paths and Codex quick start, and the QCode availability date from the OpenAI DevDay 2026 page on qcode.cc, all crawled the same day.
Sol with half-price cached input, ready to call
Call gpt-6.1-sol, gpt-6-sol and gpt-6-astra from one endpoint, billed per token; prices are on /models.
Related reading
OpenAI DevDay 2026 announcement recap
OpenAI's own sentences from 2026-09-29, item by item: gpt-6.1-sol, three API changes and what is still unpublished.
GPT-6 Sol / Luna pricing tracker
Official gpt-6-sol rates, billing rules and recomputable examples.
GPT-6 Astra complete guide
The top GPT-6 tier: specs, pricing and four migration constraints.
The figures and quotes on this page were checked on 2026-10-07 against OpenAI's official model pages, price list and changelog; OpenAI may change them without notice. Claims such as "near-Astra" are OpenAI's own statements and testing, not a third-party audit, and are not a promise about your specific workload. Worked examples are converted from the official price list; QCode bills per token, with per-model prices on /models. Model availability is whatever /models shows.