Claude Code's 1M Context Behind a Custom Base URL: What to Do When Your Gateway Stops at 200K
As of 2026-10-07: since Claude Code 2.1.285 (released 2026-09-29), setting a custom ANTHROPIC_BASE_URL makes Opus 4.7 and later, Sonnet 5 and later and the Fable models run with a 1M context window automatically; if your gateway stops at 200K, the official advice is to run /autocompact 200k. The same run of releases added the x-claude-code-prompt-id gateway hint header and the CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS backoff variable for 529 retries. This page covers all three and the exact settings.
Updated 2026-10-07
Four numbers to know first
Default window behind a custom base URL
From 2.1.285, Opus 4.7 and later, Sonnet 5 and later and the Fable models run at 1M behind a gateway too, with no [1m] variant to pick.
When your gateway stops at 200K
The changelog says to run /autocompact 200k if your gateway stops at 200K; Claude Code cannot detect a lower limit that the gateway enforces.
Range /autocompact accepts
The command takes forms like 200k or 1M; the CLAUDE_CODE_AUTO_COMPACT_WINDOW variable takes a plain integer only, and 500k reads as 500 and clamps to the 100K minimum.
Overload retry backoff (2.1.292)
The new CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS variable sets a longer base delay for the backoff when retrying a 529 overloaded request.
What happens once ANTHROPIC_BASE_URL is set
As of 2026-10-07, the answer is: behind a custom ANTHROPIC_BASE_URL, Claude Code gives each model it recognizes the same context window the model has on the Anthropic API. The docs state that Fable 5.1, Fable 5, Sonnet 5 and later, and Opus 4.7 and later get the 1M window with no [1m] variant to select, while a model that reaches 1M only through its [1m] variant (the docs name Opus 4.6; Sonnet 4.6 works the same way) runs at 200K without it. The key sentence is "Claude Code can't detect a lower limit that the gateway or the server behind it enforces": if the gateway or the server behind it caps context lower, the client does not know and keeps budgeting for 1M until requests are rejected.
What changed in this run of releases (2026-09-25 to 2026-10-06)
In order of GitHub Releases publish time (UTC): 2.1.283 (2026-09-25) added x-claude-code-prompt-id to the gateway hint headers, opt in with CLAUDE_CODE_GATEWAY_HINT_HEADERS=1; 2.1.285 (2026-09-29) switched sessions behind a custom ANTHROPIC_BASE_URL to the 1M window of models that have one; 2.1.286 (2026-09-30) changed retries so one limit covers a whole model call, meaning a failing call sends at most 14 requests with default settings, and made the fallback notice say when a fallback dropped the window from 1M to 200K; 2.1.288 (2026-10-02) made /autocompact save the window per model; 2.1.292 (2026-10-06) added CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS to set a longer base delay for 529 retries.
Three key releases
Claude Code 2.1.283 ships: the gateway hint headers gain x-claude-code-prompt-id, so a gateway can group the requests that serve one user prompt.
Claude Code 2.1.285 ships: sessions behind a custom ANTHROPIC_BASE_URL use the model's own 1M window; run /autocompact 200k if your gateway stops at 200K.
Claude Code 2.1.292 ships with CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS for a longer 529 retry base delay; this page was checked against the official pages on 2026-10-07.
Confirmed vs not verified
Confirmed (verbatim in the official pages)
All of the following can be checked word for word in the Claude Code changelog and official docs: from 2.1.285, sessions behind a custom ANTHROPIC_BASE_URL use each model's own window (1M for Opus 4.7+, Sonnet 5+ and Fable); the official advice is /autocompact 200k when a gateway stops at 200K; Claude Code cannot detect a lower gateway limit; since 2.1.288 /autocompact saves per model; the command accepts 100K to 1M; CLAUDE_CODE_AUTO_COMPACT_WINDOW takes a plain integer only and overrides the command, the flag and the setting; CLAUDE_CODE_DISABLE_1M_CONTEXT=1 turns 1M off; x-claude-code-prompt-id exists from 2.1.283 and is off by default behind a custom base URL; CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS exists from 2.1.292.
Unverified and not published by the vendor
Three things are not verified, so don't decide based on them: first, CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS has only a one-line changelog entry so far, with no published default or ceiling; second, how much context a specific gateway or relay (QCode included) actually accepts for each model is not something the official docs answer or Claude Code can detect, so go by your own tests; third, whether a gateway reads x-claude-code-prompt-id and what it does with it is up to each gateway, and the docs only say it can be used to group a prompt's requests.
Two ways to pull the window back, and how to choose
Lower the auto-compact threshold (/autocompact or the variable)
Changes only when compaction runs, not the model picker. Since 2.1.288, /autocompact 200k is saved for the current model only, so set it again after switching; to cover every model, put "autoCompactWindow": 200000 in ~/.claude/settings.json, or set CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000 in the environment that starts Claude Code, which takes precedence over the rest. The floor is 100K, so for a gateway limit below that, /compact remains the recovery.
Turn 1M off entirely (CLAUDE_CODE_DISABLE_1M_CONTEXT=1)
Suits a team-wide cap. 1M variants disappear from the model picker, and native-1M models such as Sonnet 5 and the Fable models are treated as having a 200K window: with auto-compaction on they compact at the 200K boundary, and with it off they stop there with the context-limit error. Setting the auto-compact window above 200K doesn't lift the hold.
Gateway stops at 200K: four steps
① Find out the gateway's limit: Claude Code cannot detect a lower limit that the gateway enforces, so check the gateway's docs or test it yourself. ② For the current model only, run /autocompact 200k in the session (saved per model since 2.1.288, so repeat it after switching; /autocompact auto returns to the window tuned for the model). ③ For every model, put "autoCompactWindow": 200000 in ~/.claude/settings.json, or export CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000 before launch. A window saved for a model with /autocompact takes precedence over autoCompactWindow in the same file; the variable beats both but accepts only a plain integer, and a suffixed value like 500k reads as 500 and clamps to the 100K minimum. ④ If you already hit the error, run /compact to recover the session. If the gateway rewrites the too-long error in its own words (the docs cite ContextWindowExceededError and prompt token count of N exceeds the limit of M), Claude Code does not recognize it and won't compact and retry automatically, so run /compact by hand.
On QCode
QCode's docs say connecting Claude Code takes two environment variables: ANTHROPIC_BASE_URL set to https://api.qcode.cc/api (no trailing slash) and ANTHROPIC_AUTH_TOKEN set to the key starting with cr_ that you create in the console. Users in North America and Europe can swap the host for us.qcode.cc (Los Angeles) and keep the same key. This page makes no promise about the context limit of any model on QCode; if you hit a context-limit error, follow Claude Code's official advice and use /autocompact to lower the threshold to what actually goes through. To see the model, context length and usage of each request, the docs point to probe.qcode.cc, where you enter your cr_ key. Billing is per token; see /models for each model's price.
Frequently asked questions
Why does Claude Code show a 1M context after I switched to a relay?
Because since 2.1.285 (released 2026-09-29), Claude Code applies the Anthropic API window behind a custom ANTHROPIC_BASE_URL too: Opus 4.7 and later, Sonnet 5 and later and the Fable models get 1M with no [1m] to pick. Opus 4.6 and Sonnet 4.6 reach 1M only with [1m] and run at 200K without it, and the official context-window page lists 200k for Claude Sonnet 4.5 and Claude Haiku 4.5. That is the client's budget, not proof that your gateway accepts 1M.
My gateway only supports 200K. What error will I see?
Requests over the limit are rejected by the gateway. The too-long error Claude Code recognizes is Prompt is too long, shown in an interactive session as Context limit reached · /compact or /clear to continue; if the gateway rewrites the error in its own words, such as ContextWindowExceededError, Claude Code does not recognize it and won't compact and retry automatically. Either way, run /compact to recover, then lower the compaction threshold to 200K as the docs advise.
Does /autocompact 200k apply to every model?
Not necessarily. Since 2.1.288 (released 2026-10-02), /autocompact saves the window per model in your user settings, so you set it again after switching; before that it saved one window for every model. To cover every model at once, put "autoCompactWindow": 200000 in ~/.claude/settings.json or set CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000. The variable takes precedence over the command, the flag and the setting, and accepts a plain integer only.
What is x-claude-code-prompt-id, and should I turn it on?
It is a random UUID added to the gateway hint headers in 2.1.283: requests serving one user prompt, including the turns of subagents that prompt started, share the value, so a gateway can group them. It is sent by default on a direct connection to the Anthropic API but off by default behind a custom base URL, because a proxy that rejects unknown headers would fail the request; to send it, set CLAUDE_CODE_GATEWAY_HINT_HEADERS=1. The docs state the hint headers never carry prompt text or file contents.
What does CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS do?
It sets a longer base delay for the backoff when Claude Code retries a 529 overloaded request; it arrived in 2.1.292 (released 2026-10-06), and no default value has been published. Separately, since 2.1.286 one retry limit covers a whole model call, so with default settings a failing call sends at most 14 requests; for unattended jobs, CLAUDE_CODE_RETRY_WATCHDOG retries 429 and 529 capacity errors indefinitely.
Can I use the full 1M with Claude Code on QCode?
This page makes no promise on that. Claude Code budgets for the window the model has on the Anthropic API, but it cannot detect a lower limit that a gateway enforces. To connect, set ANTHROPIC_BASE_URL to https://api.qcode.cc/api and ANTHROPIC_AUTH_TOKEN to your cr_ key; if you hit a context-limit error, run /compact to recover the session, then follow the official advice and use /autocompact to lower the threshold to what actually goes through.
Sources
The Claude Code changelog (CHANGELOG.md in the anthropics/claude-code GitHub repository, fetched 2026-10-07): the entries for 2.1.283, 2.1.285, 2.1.286, 2.1.288 and 2.1.292, with release dates taken from the publish times (UTC) in the same repository's GitHub Releases. The official Claude Code docs at code.claude.com (fetched 2026-10-07): model configuration (the context window behind a gateway, the auto-compact window, the 1M switch), environment variables, connecting to an LLM gateway (the troubleshooting row on rewritten too-long errors), the gateway compatibility guide (hint headers) and the error reference. Anthropic's platform docs page on context windows (fetched 2026-10-07). QCode setup: the environment variables, endpoints and API formats, and Claude Code feature availability pages on docs.qcode.cc (fetched 2026-10-07).
Connect Claude Code to QCode
Two environment variables, one cr_ key for Claude models, billed per token; see /models for each model's price.
Related reading
Context Length Exceeded troubleshooting
The three disguises of context overflow and how to tell them apart.
Claude Code custom endpoint setup
The two environment variables, the settings.json form, apiKeyHelper, and the documented web-session limitation.
How to code efficiently inside 1M+
Repo maps, cache prefixes, selective loading and anti-patterns.
Checked on 2026-10-07 against the official Anthropic and Claude Code pages, which take precedence; Claude Code ships every week, so behavior may change between versions without notice. This page makes no promise about context limits on QCode; model availability is whatever /models shows.