Cognition, the company behind the Devin autonomous coding agent, released SWE-2 on September 10, its newest agentic coding model. The company says SWE-2 scored 50.0% on FrontierCode 1.1 Main, a benchmark within one point of the score Cognition attributes to a model it calls Fable 5.1, while claiming SWE-2 costs 64% less to run per task.

Every number in Cognition’s announcement is Cognition’s own measurement, run on a benchmark and cost basis the company chose. The company did not publish independent verification, and its cost comparisons use list pricing that includes public discounts, a detail buried in an appendix rather than the main post.

SWE-2 is a post-trained version of Kimi K3, a 2.8 trillion parameter open model, built on infrastructure Cognition first used for its earlier SWE-1.7 release. Cognition says it scaled reinforcement learning to that parameter count for the first time.

On the company’s own numbers, SWE-2 outperforms both SWE-1.7 (its predecessor) and a model it calls Grok 4.6 on score and cost together. It says SWE-2 ties a model it labels GPT-5.6 Sol, along with two versions of Fable, at a lower price, and lands a few points behind a model called GPT-6 Astra while costing roughly a quarter as much. On a separate benchmark, Terminal-Bench 4, Cognition’s own figures show SWE-2 trailing both Fable 5.1 and GPT-6 Astra by wide margins, a gap the post does not explain.

Cognition frames the efficiency gains as much a workflow story as a scoring one. The company says its medium-effort setting outperforms SWE-1.7 while taking 58% fewer conversational turns and costing 81% less, and that the model begins editing code after a median of 18 exploration steps, versus 48 for its predecessor.

Cognition also revisited its own trustworthiness testing, an evaluation it introduced earlier this year. On a 145-question set covering topics China treats as politically sensitive, run in English and two forms of Chinese, the company reports SWE-2 passing 98.0% of attempts overall, with the lowest pass rate, 95.2%, in Simplified Chinese. Cognition tested SWE-2 against Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1, and a model called Opus 5, using its own binary judge rather than an outside evaluator.

The model is live now in Devin Desktop and the Devin command line tool, with a rollout to Devin Web and Fusion still in progress.

The real test for SWE-2 is not Cognition’s chart. It is whether Devin’s paying customers see the same 64% cost reduction on their own workloads, since list-price comparisons and internally chosen benchmarks routinely overstate real-world savings once discounts, retries, and edge cases enter the picture. Any team evaluating Devin against a rival agent this quarter should ask Cognition for cost data on its own task mix before treating the FrontierCode numbers as representative.

Cognition published these findings on its company blog on September 10, 2026.