Kimi K3 vs Opus 5
Overlapping scores, different products
Third-party aggregates put Opus 5 slightly ahead overall, with ranges that overlap K3's; K3's advantages are 2.8T open weights, roughly 40% lower cost and native vision. Moonshot itself acknowledges the model still trails Fable 5 and GPT-5.6 Sol.
Highlights
K3 scale
16/896 MoE, KDA+AttnRes, native vision, 1M.
Cheaper tokens
K3 $3/$15 vs commonly cited Opus 5 $5/$25 (verify).
Aggregates
BenchLM-style pages ~83 vs ~80, intervals overlap.
K3 blog
Still trails Fable 5 and GPT-5.6 Sol overall.
The matchup
K3's headline demos — kernel optimisation, MiniTriton, chip design, game generation — are capability showcases Moonshot published itself. They illustrate what K3 does well; they are not evidence that Opus 5 loses on those tasks, since no same-source head-to-head accompanies them.
Why it is trending
Limitations from the K3 post: must return full thinking history; can be overly proactive; UX still behind Fable/Sol. Those belong on this page.
Timeline
K3 GA + blog.
Weights promised (custom commercial clause — verify license).
Third-party vs Opus 5 writeups.
Confirmed vs caution
Confirmed
Official K3 specs/price/limitations; third-party aggregates with overlap; Opus 5 as closed flagship.
Caution
The two overlap, and neither dominates. Moonshot's own blog post states plainly that K3 still trails Fable 5 and GPT-5.6 Sol overall. On price, Claude Opus 5 lists officially at $5/M input and $25/M output, against Kimi K3's $3/M and $15/M.
How to choose
Pick Opus 5
Highest independent aggregates, Claude stack, speed-sensitive coding (some reviews ~2×).
Pick K3
Open 2.8T, vision+code, $3/$15, self-host if you can feed a supernode.
How to use it
K3 via Kimi API/Work/Code. Opus 5 via existing Claude paths. Cross-link Work + Opus 5 guide.
On QCode
QCode currently lists both kimi-k3 and claude-opus-5 on /models, so you can compare them with a single key.
FAQ
Did K3 beat Opus 5?
No clean win. Overlap on aggregates; Opus usually ahead.
Price gap?
About 40% on the $3/$15 vs $5/$25 quote — verify.
Local K3?
2.8T. Official serving hint: 64+ accelerator supernode. Not a 16G story.
Thinking history?
K3 wants full thinking replay. Mid-session model swaps can get ugly.
Vs Fable 5 instead?
In its own K3 blog post, Moonshot states that overall capability still trails Claude Fable 5 and GPT-5.6 Sol, while leading other open models on a number of measures. Reading third-party leaderboards with that sentence in mind gives a much more accurate picture.
DeepSeek 0813 in this fight?
Yes. DeepSeek's same-source 0813 grid on Hugging Face includes a Kimi K3 row too (HLE 43.5/56.0, Terminal-Bench 2.1 88.3, DeepSWE 67.5, Toolathlon 76.5); the full table is on /deepseek-v4-pro-0813.
Sources
kimi.com/blog/kimi-k3 ; BenchLM/CodingFleet-style compares (cite as third party); Opus 5 site pages.
Disruptive, not the aggregate king
That sentence is the honest H2.
Related
Kimi Work
How to run long K3 jobs.
K3 tracker
K3's release timeline and open-weights progress.
Opus 5 guide
Closed-side hero page.
Not affiliated with Moonshot or Anthropic.