Cognition SWE-2: the vendor table, the cost claims, and what is missing
Cognition published SWE-2 on 2026-09-10, calling it their most advanced coding model, post-trained from Kimi K3 (2.8T parameters) and aimed at the Pareto frontier of capability versus inference cost: 50.0% on FrontierCode 1.1 Main, within a point of Fable 5.1 while being 64% cheaper. The same table shows 27.3% on Terminal-Bench 4. This page separates what the vendor states, what it does not, and what you can run here.
Updated 2026-09-20
Four numbers from the official table
FrontierCode 1.1 Main
Their words: achieving 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper. The same row lists SWE-1.7 42.0%, Kimi K3 44.2%, Grok 4.6 48.0%, Fable 5.1 50.9%, GPT-6 Astra 53.3% - all vendor-reported.
Claimed reduction versus Fable 5.1
The post says being 64% cheaper. That is Cognition's own comparison: no third-party re-run and no published price per million tokens behind it.
Terminal-Bench 4 (the weak column in the same table)
Same row: SWE-2 27.3%, Kimi K3 21.5%, Grok 4.6 20.3%, Fable 5.1 55.8%, GPT-5.6 Sol 37.3%, GPT-6 Astra 57.9%, SWE-1.7 7.6%. Cognition published this line itself; we are not omitting it.
Base-model parameters (Moonshot's figure)
Original: SWE-2 is post-trained from Kimi K3, a 2.8T-parameter model. 2.8T describes the base K3, not SWE-2's total size; Cognition separately says it scaled RL to the multi-trillion-parameter regime.
What SWE-2 is
SWE-2 is the coding model Cognition (the Devin company) released on 2026-09-10, described as their most advanced coding model yet, with competitive agentic coding across multiple effort levels. It is not trained from scratch: the post says it is post-trained from Kimi K3, Moonshot's open-weight model, with RL that the company claims adds 5-6 points on many benchmarks. It shipped in Devin Desktop and Devin CLI the same day, rolling out on Devin Web and Fusion. The route is product subscription plus agent orchestration - the post lists no public endpoint, no price per million tokens and no rate limit.
Release cadence
2026-09-10: Cognition publishes Introducing SWE-2: Pushing the Pareto Frontier; the page metadata dates it 2026-09-10 and the body says available starting today in Devin Desktop and CLI, with Web and Fusion to follow. This page was checked on 2026-09-20: that post is still the only source for the table, and no third-party reproduction exists.
Three dated entries
2026-09-10: Cognition releases SWE-2. The official table gives 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, 27.3% on Terminal-Bench 4, and claims 64% lower cost than Fable 5.1.
Same post, 2026-09-10: SWE-2 is post-trained from Kimi K3 (2.8T) on the SWE-1.7 infrastructure and recipe, with RL the company says still finds substantial headroom, adding 5-6 points.
2026-09-20, checked for this page: searching the public QCode model catalog for SWE-2 / Devin / Cognition returns no match, so nothing here claims the model is callable, and the coding links below point at models we actually sell.
Stated versus not yet sayable
What the post states
As of 2026-09-20 that page still says: the four benchmark scores (50.0% / 73.0% / 92.8% / 27.3%), 64% cheaper than Fable 5.1, base model Kimi K3 (2.8T), available the same day in Devin Desktop and CLI, scored across multiple effort levels. All of it is the vendor stating things, not an independent measure.
Do not read these as facts
As of 2026-09-20 these are absent: a public API endpoint and model id, price per million tokens, rate limits, context length, training data and licence, whether weights are open, enterprise terms. A claim circulating in search results about a free month on Pro/Max/Teams does not appear anywhere in the fetched post, so this page does not repeat it. Within a few points of GPT-6 Astra at a quarter of the cost is the vendor's own comparison, and every score comes from Cognition's table with no third-party re-run.
SWE-2 versus running your own coding agent
Cognition SWE-2 (inside the product)
Model plus Devin's agent orchestration in one product, available in Desktop and CLI from day one. The trade-off is subscription rather than per-token billing: with no endpoint, unit price or rate limit published, cost per task can only be computed on their terms.
Sold general models plus your own harness
Run coding tasks on models QCode sells: swap the model id, pay per token, keep rate limits and invoices in your own console. The trade-off is that orchestration, context management and the verification loop are yours to build, and effort levels have to be approximated.
Three things if you want the same shape
First, task splitting: the vendor scores SWE-2 across effort levels; a self-built stack has no such internal tiers, so teams split locate / patch / test into separate calls and route simple steps to cheaper models. Second, the verification loop: the most concrete sentence in the post is about training a model that iteratively hardens our verifiers - in engineering terms tests and lint must be automated or any model's score collapses. Third, cost claims: 64% is a vendor comparison, not a unit price. Benchmark it on your own repositories and issue set, record tokens and retries, then read your invoice.
What is actually available on QCode
SWE-2 is not offered. On 2026-09-20 the public QCode model catalog returns 0 matches for SWE-2 / Devin / Cognition, and the 30-day priced-usage table lists none of them either, so this page states no availability and no price. To land coding-agent work on billable calls, the route is sold general models plus schema/tool calling with outer validation and retries, billed as official price times service rate; unit prices and rate limits live on /pricing and in the live catalog.
Questions
Can I call SWE-2 through QCode?
No. The 2026-09-20 catalog check finds no SWE-2 and no Cognition or Devin service; it is only served inside the Devin product line (Desktop, CLI, Web and Fusion rolling out).
Are those benchmark numbers trustworthy?
They are vendor-reported: the five scores and the 64% figure all come from the same Cognition post, and no third-party re-run was published as of 2026-09-20. To their credit the weak column is there too: 27.3% on Terminal-Bench 4.
What is its relationship to Kimi K3?
The post says SWE-2 is post-trained from Kimi K3 (2.8T), so the base is Moonshot's open-weight model and Cognition adds RL on top. K3's own row is different: 44.2% on FrontierCode 1.1 Main and 21.5% on Terminal-Bench 4 in the same table.
Is there an API, a price, a context length?
The post gives none: no public endpoint or model id, no price per million tokens, no rate limit, no context length. This page does not infer them; teams needing those numbers must wait for the vendor or evaluate the product subscription.
So which route for my own coding agent?
Depends whether you want a product that runs the loop for you, or billable calls where you control model choice and context. The second lands today on sold general models plus tool calling: change the model id, pay per token, keep logs and limits in your console.
Can I use the 64% cheaper number directly?
Not as a unit price. It is Cognition comparing its own cost to Fable 5.1, with no published pricing and no third-party recalculation, and the same sentence carries within a few points of GPT-6 Astra at a quarter of the cost - also vendor framing. Measure on your own tasks.
Sources
Cognition, Introducing SWE-2: Pushing the Pareto Frontier, 2026-09-10, https://cognition.com/blog/swe-2 (fetched 2026-09-20); the QCode catalog and usage figures are our own internal check (public prod catalog snapshot plus the 30-day priced-usage table). Every score and percentage here is vendor-reported.
Skip the agent product and still run coding work on billable calls
SWE-2 is not in our catalog (checked 2026-09-20); what you can use is sold general models plus tool calling, billed per token.
Read next
Claude Fable 5 on SWE-Bench Pro
The comparator Cognition measures against: where the number comes from and its limits.
Choosing AI coding models, H2 2026
Budget, complexity and context for building a coding agent on sold models.
QCode pricing guide
Official price times service rate, plus how caching and retries are billed.
Unaffiliated with Cognition or Moonshot AI. Every score, percentage and training description here paraphrases the Cognition post fetched on 2026-09-20 and is vendor-reported; QCode does not offer SWE-2 and does not evaluate the Devin product.