Claude Adaptive Thinking and the Effort Parameter, Model by Model
As of 2026-10-08, every Claude 5.x model has adaptive thinking on by default: on Opus 5.5, Fable 5.1 and Fable 5 thinking cannot be turned off; Sonnet 5.5 rejects disabled and its lowest setting is between_tools; Opus 5 and Haiku 5.5 accept disabled only at high effort or below; Sonnet 5 can turn it off outright. Thinking depth is now set with output_config.effort (low / medium / high / xhigh / max), which defaults to medium on Opus 5.5 and Haiku 5.5 and high elsewhere, and budget_tokens returns a 400 error on Claude 4.7 and every later model.
Updated 2026-10-08
Four conclusions to remember
Thinking on Opus 5.5 / Fable 5.1
The official docs say Opus 5.5, Fable 5.1 and Fable 5 reject thinking disabled, so thinking cannot be turned off; on Opus 5.5, sending disabled returns a 400 error at every effort level.
Default effort on the API
Officially, most models default to high, while Opus 5.5 and Haiku 5.5 default to medium. Passing the default explicitly is the same as omitting it and does not break the prompt cache.
Effort values
low, medium, high, xhigh and max. Opus 5.5, Sonnet 5.5 and Haiku 5.5 support all five; the docs note that not every model that supports max supports xhigh.
Where budget_tokens stands
The docs state that Claude 4.7 and later models reject type enabled with budget_tokens; on the Claude 4.6 models it still succeeds but is deprecated; on 4.5 and earlier it is the only manual mode available.
What adaptive thinking and effort each control
As of 2026-10-08, the official definitions are: with adaptive thinking the model evaluates each request and determines whether to think and how much, and the effort parameter controls how many tokens Claude spends when responding to requests. The two do different jobs: thinking decides whether Claude reasons in thinking blocks first, while effort sets how much work goes into the whole response, including thinking, text and tool calls. The docs warn that adaptive is a thinking mode, not an effort level, so never pass it as an effort value. Effort lives at output_config.effort, not inside the thinking object, and it is a behavioral signal rather than a strict token budget, so no level guarantees a thinking block on every request.
Recent official changes
Opus 5.5, released 2026-09-22, lowered the default effort to medium: the docs note that Opus 5 and earlier Opus models default to high, so a request that omits effort runs one level lower on Opus 5.5 than on Opus 5, and its thinking is always on and cannot be turned off. Sonnet 5.5, released 2026-09-28, rejects disabled; to turn off up-front thinking you send between_tools, which works only at high effort or below. Since 2026-10-05 the Models API reports capabilities.thinking.types.disabled, so you can check before sending a request whether a model accepts disabled (supported is false when it rejects it). Haiku 5.5, released 2026-10-07, thinks by default, returns a 400 error for manual budget_tokens, and a response can begin with thinking blocks.
Timeline
Claude Opus 5.5 (claude-opus-5-5) launches: thinking is always on, both thinking disabled and enabled return a 400 error, and the default effort is medium.
Claude Sonnet 5.5 (claude-sonnet-5-5) launches: turning off up-front thinking now means sending between_tools, accepted only at high effort or below.
Claude Haiku 5.5 (claude-haiku-5-5) launches: adaptive thinking is on by default, budget_tokens returns a 400 error, and disabled turns thinking off at high effort or below.
Confirmed vs not stated by the vendor
Confirmed (verbatim in the official docs)
All of the following can be checked word for word in Anthropic's and Claude Code's official docs: thinking cannot be turned off on Opus 5.5, Fable 5.1 or Fable 5; Sonnet 5.5 rejects disabled and accepts between_tools only at high effort or below; Opus 5 and Haiku 5.5 accept disabled only at high or below; Sonnet 5 can turn thinking off; Opus 4.8, Opus 4.7, Opus 4.6 and Sonnet 4.6 do not think until you set adaptive; the official description of each effort level and each model's default; budget_tokens returns 400 on Claude 4.7 and later; thinking tokens are billed as output tokens and count toward max_tokens; and in Claude Code, thinking cannot be turned off on Opus 5.5, Sonnet 5.5, Haiku 5.5 or the Fable models.
What the vendor does not state
Three things are not published, so don't budget around them: first, how many more or fewer tokens, or how much money, each effort level costs; the docs only say effort is a behavioral signal, not a strict budget, and give no multipliers. Second, the same level name is not equivalent across models; the vendor says the effort scale is calibrated per model, and Sonnet 5.5's levels were recalibrated, so carrying old settings across models has no basis. Third, there is no official conversion of the form one effort level equals so many budget_tokens; the official migration note only says to remove budget_tokens and use effort.
How to choose: effort, max_tokens or turning thinking off
effort vs max_tokens
In the docs' own words, effort is soft guidance and max_tokens is a strict limit: max_tokens caps the total output of a request, thinking and text combined, while effort only shapes how much of that output goes to thinking and guarantees no token count. To cut cost or latency, the docs say to lower effort first, which scales the whole response down, thinking included; for a hard ceiling, use max_tokens. At high effort and above, Claude may think extensively and is more likely to exhaust max_tokens.
Lower effort vs disable thinking
Models that can turn thinking off (Sonnet 5, plus Opus 5 and Haiku 5.5 at high or below) accept disabled, but for Haiku 5.5 the docs say the better way to trade quality against speed and cost is effort. Opus 5.5 and Fable 5.1 cannot turn it off, so effort is the only lever, and on Sonnet 5.5 the lowest setting is between_tools. Changing the thinking configuration or the top-level effort starts the prompt cache over; on Opus 5.5, Sonnet 5.5, Haiku 5.5 and some other models you can change effort per message instead (beta, header mid-conversation-output-config-2026-07-01) and keep the cache. With between_tools on Sonnet 5.5 or disabled on Haiku 5.5, though, a per-message effort change returns a 400 error.
Three steps from budget_tokens to effort
First, remove budget_tokens: the official migration note says that on Claude 4.7 and later models, such as Opus 5.5, Sonnet 5, Sonnet 5.5, Fable 5.1 and Haiku 5.5, type enabled returns a 400 error, so omit thinking or send thinking type adaptive (on Opus 4.8 and Opus 4.7, omitting thinking means no thinking, so send adaptive explicitly). Second, control depth with output_config.effort: when omitted it is medium on Opus 5.5 and Haiku 5.5 and high on other models, and the docs recommend running a fresh effort sweep on your own evals rather than carrying settings over from an earlier model. Third, raise max_tokens and select content blocks by type: thinking counts toward max_tokens and a response can begin with thinking blocks; on these models display defaults to omitted (an empty thinking field with only a signature), so set display to summarized if you want the summary. usage.output_tokens_details.thinking_tokens in the response shows how many billed output tokens went to reasoning.
On QCode
According to the QCode docs, Claude models only work on the Anthropic-protocol endpoint: in Claude Code, set ANTHROPIC_BASE_URL to https://api.qcode.cc/api (the SDK builds /api/v1/messages from it). The Claude 5.x models you can call on QCode are claude-sonnet-5-5, claude-opus-5-5, claude-sonnet-5, claude-opus-5 and claude-fable-5-1, and one key switches between them by changing the model field. Whether Haiku 5.5 (claude-haiku-5-5) is listed is up to the /models page. Billing is per token, with each model's price on /models.
Frequently asked questions
Which Claude models cannot turn thinking off?
Opus 5.5, Fable 5.1 and Fable 5. The docs say these models reject thinking type disabled, in the words “Thinking can't be turned off on these models.” Sonnet 5.5 rejects disabled too, but between_tools turns off its up-front thinking; Opus 5 and Haiku 5.5 accept disabled only at high effort or below; Sonnet 5 can turn it off directly. To check before sending, read capabilities.thinking.types.disabled from the Models API.
What effort levels are there, and which is the default?
Five: low, medium, high, xhigh and max. The docs say most models default to high, Opus 5.5 and Haiku 5.5 default to medium, and Sonnet 5.5 defaults to high on the Claude API. The official descriptions: max is absolute maximum capability with no constraints on token spending, xhigh is extended capability for long-horizon work, high spends as many tokens as the task needs, medium is a balanced approach with moderate token savings, and low is the most efficient, with some capability reduction.
How much does lowering effort save?
The vendor gives no multiplier. The docs call effort a behavioral signal, not a strict token budget, that applies to every output token, including text, tool calls and thinking; low brings significant token savings with some capability reduction. Thinking tokens are billed as output tokens even when the thinking text is not returned to you; to see what went to reasoning, read usage.output_tokens_details.thinking_tokens in the response. Only max_tokens sets a hard ceiling.
Can I still use budget_tokens?
Not on Claude 4.7 and later models, where type enabled with budget_tokens returns a 400 error. On the Claude 4.6 models it still succeeds but is deprecated. Models that support only manual extended thinking, such as Sonnet 4.5, Opus 4.5 and Haiku 4.5, keep using budget_tokens; Opus 4.5 is the one such model that also supports effort, and the two combine. To migrate, remove budget_tokens, switch to adaptive and control depth with effort.
How does Haiku 5.5 differ from other 5.x models on thinking?
Per the official configuration table, Haiku 5.5 is the only one of the three 5.5 models that accepts thinking disabled: it turns thinking off at high effort or below and returns a 400 error at xhigh or max. Like Opus 5.5 it defaults to medium effort and thinks by default, and its thinking text is omitted by default, whereas Haiku 4.5 returned summarized thinking. The docs still recommend effort, rather than disabling thinking, to trade quality against speed and cost.
How do I set effort in Claude Code, and can I turn thinking off there?
Use /effort (no argument opens a slider, a level name sets it directly, /effort auto clears the saved level for the active model), the --effort launch flag, the CLAUDE_CODE_EFFORT_LEVEL environment variable, or the left and right arrow keys in /model. In Claude Code, Opus 5.5, Sonnet 5.5 and Haiku 5.5 default to medium; settings files do not accept max, and unless set through the environment variable, max applies to the current session only. The docs state that thinking cannot be turned off in Claude Code on Opus 5.5, Sonnet 5.5, Haiku 5.5 or the Fable models, and MAX_THINKING_TOKENS=0 has no effect there; writing ultrathink in a prompt only adds an in-context instruction, and the effort sent to the API is unchanged.
Sources
Anthropic's official docs (platform.claude.com, crawled 2026-10-08): the Thinking overview (per-model configuration table, rules for turning thinking off, display defaults), Troubleshooting thinking (supported and rejected configurations per model), Steering thinking (what each effort level does to thinking, cost control and billing), Effort (the five levels, per-model defaults and recommendations), Extended thinking (the deprecation of budget_tokens and the migration), the Claude Haiku 5.5 what's-new page and migration guide, the models overview, and the Claude Platform release notes (dates for Opus 5.5, Sonnet 5.5, Haiku 5.5 and the new Models API field). The Claude Code Model configuration docs (code.claude.com, same day). The QCode docs page on endpoints and API paths (docs.qcode.cc, same day).
Call Claude 5.x with one key
Set Claude Code's ANTHROPIC_BASE_URL to https://api.qcode.cc/api and put claude-sonnet-5-5, claude-opus-5-5 or another model in the model field; billed per token, with each model's price on /models.
Related reading
Claude ultracode vs max
Ultracode is a separate Claude Code toggle, not an effort level: how to turn it on and when to pick max.
Claude Sonnet 5.5 migration: the disabled-thinking error and between_tools
The exact disabled error, the limits of between_tools and effort, and computer_20251124.
The complete Claude Sonnet 5.5 guide
Specs, price list, default effort and the five breaking changes.
This page was checked on 2026-10-08 against the live versions of Anthropic's and Claude Code's official docs; upstream changes may happen without notice. How effort affects token usage varies by task and the vendor publishes no fixed multiplier, so nothing here is a promise about your specific workload. Model availability is whatever /models shows.