Claude Code with DeepSeek, GLM, Kimi and Qwen: One-Key Setup
As of 2026-10-07, using Chinese models in Claude Code through QCode works like this: keep ANTHROPIC_BASE_URL at https://api.qcode.cc/api and ANTHROPIC_AUTH_TOKEN at your cr_ key, then type /model glm-5.3 in a session (or deepseek-v4-pro, kimi-k3, qwen3.8-max), or set ANTHROPIC_MODEL before launch. QCode's docs state that the GLM, Kimi, DeepSeek and Qwen families run on both the Anthropic endpoint and the OpenAI Chat endpoint. This page sets out the setup steps, what each model suits, how they differ from Claude and common errors, based on QCode's docs and each vendor's official documentation.
Updated 2026-10-08
Four things to know first
Chinese model families on the same key
Zhipu GLM, Moonshot Kimi, DeepSeek and Alibaba's Qwen share one cr_ key and the same endpoints with Claude, per QCode's docs.
ANTHROPIC_BASE_URL stays as it is
Once Claude Code points at https://api.qcode.cc/api, switching to a Chinese model needs no address change; leave off the trailing slash.
One command inside the session
Type the ID character for character, such as deepseek-v4-pro, kimi-k3 or qwen3.8-max, not a display name like “GLM-5.3”.
GLM-5.3 context / max output (official)
Zhipu says GLM-5.3 currently handles text only and always thinks; Moonshot says kimi-k3 has a 1-million-token context window and native vision.
How it works: you change the model, not the address
As of 2026-10-07, using a Chinese model in Claude Code needs no new tool, endpoint or key: QCode's Anthropic endpoint, https://api.qcode.cc/api, accepts model IDs from Claude and from the four Chinese families, so Claude Code stays pointed where it is and /model switches it to glm-5.3, deepseek-v4-pro, kimi-k3 or qwen3.8-max. The Claude Code docs put it plainly: ANTHROPIC_BASE_URL changes where requests are sent, not which model answers them, so the model name is what you change. QCode's docs still recommend the Claude 5 family as the everyday default and describe the Chinese tier as a fit for Chinese-language office work and long documents, high-volume jobs at a lower unit price, and side-by-side checks against Claude. DeepSeek, Zhipu, Moonshot and Alibaba Cloud Model Studio each publish an official guide for pointing Claude Code at their own Anthropic-compatible endpoint; that route means a separate key and address per vendor.
Recent changes on the vendor side
DeepSeek's official news page lists DeepSeek-V4.1-Flash as released on 2026-09-10, and the company says it has native multimodal vision. DeepSeek's pricing page adds that the legacy name deepseek-v4-flash can still be called, but the model behind it has been retired and those requests are served by DeepSeek-V4.1-Flash. Zhipu's docs state that GLM-5.3 currently handles text only, supports a 1M context window and 128K maximum output, and always has thinking on, with no way to disable it. Moonshot's model list describes kimi-k3 as natively supporting vision with a 1-million-token context window. Alibaba Cloud Model Studio's Anthropic-compatible Messages docs list qwen3.8-max, qwen3.8-flash and qwen3.7-plus among supported models. These are each vendor's statements about its own platform; which ID to use on QCode follows QCode's docs and /models.
Timeline
DeepSeek-V4.1-Flash is released, per DeepSeek's official news page.
QCode's docs cross-check the Chinese IDs on sale against qcode.cc/models and the public models endpoint: glm-5.3, glm-5.3-flash, kimi-k3, deepseek-v4-pro, deepseek-v4-flash, deepseek-v4.1-flash, qwen3.8-max, qwen3.8-flash, qwen3.7-plus.
This page is checked against QCode's docs (the Chinese models page and the endpoints page were both updated 2026-09-25) and the official docs of the four vendors.
Confirmed vs not stated
Confirmed (verbatim in official pages)
Each of these can be checked word for word: QCode's docs say the four Chinese families (GLM, Kimi, DeepSeek, Qwen) can use both the Anthropic endpoint and the OpenAI Chat endpoint, while Claude models can only use the Anthropic endpoint; once Claude Code points at QCode there is no need to change ANTHROPIC_BASE_URL, and /model plus an ID switches models; ANTHROPIC_MODEL is forced on every launch, while ANTHROPIC_DEFAULT_MODEL only sets the default for new sessions (from 2.1.236); Chinese models do not appear in the /model picker even with gateway model discovery turned on; non-Claude models may not get native features such as Extended Thinking in Claude Code. On the vendor side, DeepSeek, Zhipu, Moonshot and Alibaba Cloud Model Studio all publish official guides for connecting Claude Code to their own Anthropic-compatible endpoints.
Not stated officially / not verified here
Three things are not documented or not verified here: first, QCode's docs give three ways to pick a Chinese model, namely /model, ANTHROPIC_MODEL and ANTHROPIC_DEFAULT_MODEL, and say nothing about how ANTHROPIC_DEFAULT_HAIKU_MODEL, CLAUDE_CODE_SUBAGENT_MODEL or the [1m] suffix behave on QCode (the vendor guides that use them target their own endpoints), nor whether Chinese models should set CLAUDE_CODE_MAX_CONTEXT_TOKENS, the official Claude Code variable that overrides the assumed context window; this page does not guess beyond that; second, which Chinese model codes best in Claude Code: this page cites no benchmarks and only relays QCode's pairing advice; third, the context and output limits of qwen3.8-max and the other Qwen models were not checked one by one here, so go by /models and Alibaba Cloud's official pages.
Compared with Claude models and vendor endpoints
vs Claude models
QCode's docs advise keeping Claude models as the main choice and other models as a supplement, since non-Claude models may not support parts of Extended Thinking in Claude Code. The vendor docs list their own differences; GLM-5.3, for example, handles text only and cannot turn thinking off. The Claude Code docs add that for a model ID it does not recognize, Claude Code compacts at the context window it assumes for that ID. For pairing, QCode suggests claude-sonnet-5 first for everyday coding, with deepseek-v4-pro or glm-5.3 for a comparison run or a tight budget.
vs each vendor's own endpoint
DeepSeek (https://api.deepseek.com/anthropic), Zhipu (https://open.bigmodel.cn/api/anthropic), Moonshot (https://api.moonshot.cn/anthropic) and Alibaba Cloud Model Studio all publish Claude Code guides, but each needs its own key, its own ANTHROPIC_BASE_URL, and every tier (main chat, background tasks, subagents) pointed at its own models: Moonshot's guide warns that configuring only some variables makes the matching scenarios fail silently. Through QCode, one cr_ key covers Claude and all four Chinese families, and switching changes only the model.
Setting up Chinese models in Claude Code: five steps
① Set two variables as QCode's environment docs show: ANTHROPIC_BASE_URL=https://api.qcode.cc/api (no trailing slash) and ANTHROPIC_AUTH_TOKEN=your cr_ key; for North America and Europe there is also us.qcode.cc (Los Angeles). ② Start claude and type /model glm-5.3 in the session, or another ID on sale such as deepseek-v4-pro, kimi-k3 or qwen3.8-max, copied character for character. ③ To start on a Chinese model every time, set ANTHROPIC_MODEL (forced on every launch); to set only the default for new sessions, use ANTHROPIC_DEFAULT_MODEL (Claude Code 2.1.236 and later; a /model choice overrides it). ④ Run /status to see which model is active. ⑤ For an OpenAI-compatible tool instead of Claude Code, use https://api.qcode.cc/openai/v1 with the same Chinese model IDs; Claude models cannot use this endpoint, so use Claude Code or another tool that speaks the Anthropic protocol.
On QCode
The Chinese model IDs you can call through QCode are glm-5.3, glm-5.3-flash, glm-5.2, kimi-k3, deepseek-v4-pro, deepseek-v4-flash, deepseek-v4.1-flash, qwen3.8-max, qwen3.8-flash and qwen3.7-plus, on the same cr_ key (created in the console) as Claude models such as claude-sonnet-5 and claude-opus-5-5. Claude Code uses the Anthropic endpoint https://api.qcode.cc/api; Chinese models in OpenAI-compatible tools use https://api.qcode.cc/openai/v1, while Claude models only work on the Anthropic endpoint; the key works across families. us.qcode.cc (Los Angeles) serves North America and Europe. Billing is per token; see /models for each model's unit price. Every request can be looked up at probe.qcode.cc by entering your key.
Frequently asked questions
Can Claude Code use DeepSeek, GLM, Kimi or Qwen?
Yes. Through QCode, keep ANTHROPIC_BASE_URL at https://api.qcode.cc/api and switch with /model deepseek-v4-pro (or glm-5.3, kimi-k3, qwen3.8-max) in the session. QCode's docs state that the GLM, Kimi, DeepSeek and Qwen families can use both the Anthropic endpoint and the OpenAI Chat endpoint, and the same cr_ key still calls Claude.
Do I need to change ANTHROPIC_BASE_URL?
No. QCode's docs say that once Claude Code points at QCode through the environment variables, ANTHROPIC_BASE_URL stays as it is. The Claude Code docs explain why: ANTHROPIC_BASE_URL changes where requests are sent, not which model answers them, so you change the model with /model or ANTHROPIC_MODEL instead.
Why don't Chinese models appear in the /model picker?
Claude Code only takes models whose names contain claude or anthropic when it reads a gateway's model list, so this is expected. Per QCode's docs, even with CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1 the picker shows only Claude IDs when connected to QCode. Type /model glm-5.3 directly, or set ANTHROPIC_MODEL or ANTHROPIC_DEFAULT_MODEL.
I entered the model name and got an error. What now?
Most often the ID is wrong. QCode's docs say the ID must be written like glm-5.3, not “GLM-5.3” or a display name, and retired names such as qwen3.7-max, kimi-k2.6 and glm-5.1 return errors. The Claude Code docs note that a value set through ANTHROPIC_MODEL is not checked up front, so a typo surfaces on the first request as “There's an issue with the selected model”. A 401 points to the key: it should start with cr_ and contain no spaces.
How do Chinese models differ from Claude inside Claude Code?
Mainly in native features. QCode's docs say non-Claude models may not get native features such as Extended Thinking in Claude Code, so treat them as a supplement when unsure. Vendor docs add specifics: GLM-5.3 is text-only and cannot disable thinking, while Moonshot says kimi-k3 supports vision natively. The Claude Code docs also say that for an unrecognized model ID it compacts at the context window it assumes; QCode's docs do not say how to adjust this for Chinese models.
Which Chinese model fits which job?
From QCode's pairing table: for everyday coding start with claude-sonnet-5, and use deepseek-v4-pro or glm-5.3 for a comparison run or a tight budget; for hard reasoning or big refactors claude-opus-5 comes first, with kimi-k3 to try if you want to save; for Chinese long-form writing, minutes and office work, kimi-k3, glm-5.3 and qwen3.8-max are common; for bulk chores, deepseek-v4-flash and qwen3.8-flash have a lower unit price.
Sources
QCode: the Chinese models page, the endpoints and API formats page and the environment variables page of QCode's docs (all updated 2026-09-25), plus the model selection guide and the billing page, crawled 2026-10-07. Claude Code: the model configuration and environment variables pages of Anthropic's official Claude Code docs (what ANTHROPIC_BASE_URL does, the precedence of /model and ANTHROPIC_MODEL, /status, the Haiku-tier variable used for background functionality, and the context window assumed for unrecognized model IDs), crawled the same day. Vendors: DeepSeek's API docs (Claude Code integration, the Anthropic API, models and pricing, the V4.1 Flash release news); Zhipu's open platform docs (Claude API compatibility, GLM-5.3, GLM-5.3-Flash and the Coding Plan model-switching page); Kimi's open platform (using Kimi in Claude Code, the model list); the Alibaba Cloud Model Studio help center (Claude Code, Anthropic-compatible Messages); all crawled 2026-10-07.
One key, Chinese models in Claude Code
Claude, GLM, Kimi, DeepSeek and Qwen share one cr_ key; switching changes only the model. Billed per token, see /models for each model's unit price.
Related reading
GLM-5.3 vs Kimi K3: routes and prices compared
Two Chinese flagships compared on access routes and price.
DeepSeek Harness + Cordis: developer preview guide
DeepSeek's official agent harness, dsh, in developer preview.
Claude Code user guide
A complete tutorial from installation to mastery.
The setup steps and quotes on this page were checked on 2026-10-07 against QCode's docs and the official pages of DeepSeek, Zhipu, Moonshot, Alibaba Cloud Model Studio and Anthropic; upstream changes may happen without notice. Endpoint and variable settings in vendor guides apply only to each vendor's own platform; model availability is whatever /models shows.