TLDRocket
Sign in

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

MarkTechPost Asif Razzaq Covered by 2 sources

Cognition launched SWE-2, its strongest coding model yet, but only inside Devin. It’s close to Fable 5.1 on FrontierCode while costing 64% less.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Cognition, the team behind the Devin coding agent, has put out SWE-2, its most capable coding model so far. The model is post-trained with reinforcement learning on Kimi K3, Moonshot AI’s 2.8T-parameter open model, and Cognition says that setup gets it to a FrontierCode 1.1 Main score of 50.0%, within a point of Fable 5.1 while cutting cost by 64%.

The catch is distribution. SWE-2 does not ship with open weights, and there’s no standalone API. It only runs inside Devin right now, through Desktop and CLI, with Devin Web and Fusion still rolling out. For a model that just posted a strong benchmark showing, that’s a pretty tight leash.

Cognition says SWE-2 is the first of its models to use selectable reasoning-effort levels, all trained in a single RL run. The company scaled that process up to a multi-trillion-parameter regime, using a base model with almost three times the parameters of its earlier setup. It also says the RL process still finds room to improve K3, adding 5 to 6 points on many benchmarks.

On the published table, SWE-2 leads Terminal-Bench 2.1 and beats its K3 base on every row. But it clearly has a weak spot: Terminal-Bench 4, where it trails Fable 5.1 and GPT-6 Astra by roughly 30 points. Cognition’s own framing is that this is a tradeoff, not a victory lap.

The more interesting change may be behavioral rather than numerical. Cognition says SWE-1.7 used to wander on simpler tasks, while SWE-2 uses what it calls focused exploration. On FrontierCode 1.1 Main, the medium setting scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less. The first real edit arrives after a median of 18 steps, compared with 48 for SWE-1.7. That is the kind of number that matters when a coding agent is supposed to feel less like a pinball machine.

My take — AI-written commentary, not fact-checked reporting

This is the usual frontier-model move: impressive numbers, then a locked door. If a model is good enough to brag about cost efficiency, it’s good enough to make the access story less theatrical. Open weights would let people test whether the gains are real outside Devin’s walled garden, but closed distribution keeps the applause on a short leash.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.