Guide · listed as kimi-k3

How to call Kimi K3
2.8T · 1M · thinking will not turn off

Moonshot shipped Kimi K3 on 16 Jul 2026. Official API: $3 cache-miss / $0.30 cache-hit / $15 out per million tokens, flat across the 1M window. Model id: kimi-k3. QCode /models lists it. The status tracker and vs Opus 5 pages already exist — this is the how-to.

#kimi-k3#2.8T#1M#$3 / $0.30 / $15

Four numbers before you call

2.8T

Scale

Official: 2.8T, native vision, KDA + Attention Residuals. They call it the first open 3T-class model, and still say it trails Fable 5 and GPT-5.6 Sol — that is their own positioning.

1M

Context

Official 1,000,000-token window. The API tariff is flat across that window. No length surcharge is written on the official table.

$3 / $0.30 / $15

Official API / million

Cache-miss $3, cache-hit $0.30, output $15. From platform.kimi.ai. Same band across the full 1M window.

Always on

You cannot disable CoT

Official FAQ: You can’t — K3 always thinks. Do not pass a thinking parameter. reasoning_effort is low / high / max, default max.

What K3 is

Moonshot’s flagship thinking model. The launch blog sells long-horizon coding, knowledge work and reasoning, with native multimodality. Launch default was max thinking; later docs added low / high / max. Thinking itself still will not switch off. Weights were promised under Modified MIT for late July — whether they are downloadable today is a Hugging Face check, not a sentence we will write as fact.

Do not mix this with the other two pages

/kimi-k3-status-tracker tracks what is confirmed. /kimi-k3-vs-claude-opus-5 is the matchup. This page is only how to call it: id, price, thinking, window, and how we talk about the weight promise. Third-party reseller rates are not the official table.

Timeline

16 Jul 2026

Kimi K3 launches on kimi.com / Kimi Work / Kimi Code / platform.kimi.ai. Id kimi-k3.

late Jul 2026

Official promise of full weights (Modified MIT). This page does not claim they are downloadable today. Check Hugging Face.

This page

How-to. QCode /models lists kimi-k3. Billing is official list × service rate.

Confirmed vs treat carefully

Confirmed

Ship date 16 Jul 2026; 2.8T; 1M context; native vision; id kimi-k3; official $3 / $0.30 / $15; thinking always on; effort low / high / max, default max. Sources: kimi.com/blog/kimi-k3 and platform.kimi.ai.

Treat carefully

Do not write “open weights today” unless you also say check HF — the promise was late July. Do not treat a reseller rate as Moonshot’s tariff. Do not invent an independent rank. Moonshot itself says it still trails Fable 5 and GPT-5.6 Sol.

Hosted API or self-host

Call the API

model=kimi-k3. Price $3 / $0.30 / $15. Thinking cannot be disabled; only reasoning_effort moves. QCode lists this id.

Want the weights

Promise was Modified MIT, late July. Whether they are up is a Hugging Face check. A blog promise is not a download URL.

How to call it

Pick kimi-k3 on the official platform. Do not pass thinking; it is not how you turn reasoning off. To shorten traces, set reasoning_effort to low. Official caveat: K3 is sensitive to thinking history — switching models mid-session or dropping prior thoughts can make quality unstable. Moonshot says coding workloads often see cache-hit rates above 90%.

On QCode

QCode /models lists kimi-k3. We bill official list × service rate. This is a how-to, not a teaser for a missing SKU. Live stock is whatever that page shows.

Kimi K3 how-to FAQ

What is the model id?

kimi-k3. That is the string on the official API platform and on QCode /models.

Can I turn chain-of-thought off?

No. Official text: You can’t — K3 always thinks. If it runs long, set reasoning_effort to low. Do not pass a thinking parameter.

What is the official price?

Per million tokens: $3 cache-miss, $0.30 cache-hit, $15 output. Flat across the 1M window. From platform.kimi.ai.

Are the weights downloadable?

They were promised under Modified MIT for late July. This page does not claim they are downloadable today. Check Hugging Face for a card.

Can I call it on QCode?

Yes. /models lists kimi-k3. Debit = official list × service rate.

How is this different from the tracker and the vs page?

The tracker is confirmation status. The vs page is Opus 5. This page is only how to call K3. Do not copy specs across them as new conclusions.

Sources

kimi.com/blog/kimi-k3; platform.kimi.ai Kimi K3 quickstart, Thinking Models, and pricing. Weight promise per the late-July blog wording; downloads per Hugging Face.

The id is kimi-k3

It is listed. Thinking stays on. Price follows the official table.

Related

Not affiliated with Moonshot / Kimi. Prices are from the official API docs. Whether weights download is a Hugging Face fact.