

Cognition's coding model that scores 50.0% on FrontierCode 1.1 Main at 64% lower cost than comparable frontier models.

Cognition's coding model that scores 50.0% on FrontierCode 1.1 Main at 64% lower cost than comparable frontier models.
SWE-2 is Cognition's most advanced coding model, post-trained from the 2.8-trillion-parameter Kimi K3 base and positioned to push the capability-versus-cost Pareto frontier rather than raw benchmark score alone. It reaches 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less, and leads on DeepSWE 1.1 (73.0%) and Terminal-Bench 2.1 (92.8%) against its own predecessor SWE-1.7 and Grok 4.6. The headline training change is a reinforcement learning algorithm that trains every reasoning-effort level in a single run, applying a linear cost penalty per effort level tuned to the local slope of the base model's frontier so the whole cost-performance curve advances while keeping its shape. Behaviorally the model explores far more selectively than SWE-1.7: on FrontierCode 1.1 Main it makes its first real edit after a median of 18 steps versus 48, finishing with 58% fewer turns and 81% lower average cost at a higher score. It is available in Devin Desktop and CLI from launch, with rollout to Devin Web and Fusion following.

Cognition's coding model that scores 50.0% on FrontierCode 1.1 Main at 64% lower cost than comparable frontier models.
SWE-2 works by combining Pareto-Frontier Cost Efficiency: Matches GPT-5.6 Sol and Fable 5/5.1 on coding benchmarks at a fraction of their price and comes within a few points of GPT-6 Astra at roughly a quarter of the cost., Single-Run Multi-Effort RL: A reinforcement learning algorithm trains all reasoning-effort levels in one run, applying a per-level linear cost penalty derived from the base model's local frontier slope., Focused Codebase Exploration: Stronger engineering judgment lets the model decide which parts of a repository matter, cutting mean steps per run from 127 to 53 at medium effort., Selectable Effort Levels: Ships medium, high and max reasoning settings so teams can trade additional steps and cost for accuracy on harder tasks., End-to-End Test Writing: Produces tests that validate an implementation end to end, catching regressions and edge cases more reliably than previous SWE models. to help users with Agentic Software Engineering: Powering Devin sessions that plan, edit, build and test changes across a real repository with minimal supervision., Cost-Sensitive Coding at Scale: Teams running large volumes of automated coding tasks pick a model that holds frontier-adjacent accuracy at a materially lower per-task cost., Terminal and Tooling Workflows: Strong Terminal-Bench results suit tasks driven through shell commands, build systems and command-line tooling., Regression Test Generation: Generating end-to-end tests for existing implementations to catch edge cases before a release., Effort-Tiered Task Routing: Routing simple tickets to medium effort and hard migrations to high or max effort within the same model deployment..
Key features include Pareto-Frontier Cost Efficiency: Matches GPT-5.6 Sol and Fable 5/5.1 on coding benchmarks at a fraction of their price and comes within a few points of GPT-6 Astra at roughly a quarter of the cost., Single-Run Multi-Effort RL: A reinforcement learning algorithm trains all reasoning-effort levels in one run, applying a per-level linear cost penalty derived from the base model's local frontier slope., Focused Codebase Exploration: Stronger engineering judgment lets the model decide which parts of a repository matter, cutting mean steps per run from 127 to 53 at medium effort., Selectable Effort Levels: Ships medium, high and max reasoning settings so teams can trade additional steps and cost for accuracy on harder tasks., End-to-End Test Writing: Produces tests that validate an implementation end to end, catching regressions and edge cases more reliably than previous SWE models..
SWE-2 is useful for anyone interested in Agentic Software Engineering: Powering Devin sessions that plan, edit, build and test changes across a real repository with minimal supervision., Cost-Sensitive Coding at Scale: Teams running large volumes of automated coding tasks pick a model that holds frontier-adjacent accuracy at a materially lower per-task cost., Terminal and Tooling Workflows: Strong Terminal-Bench results suit tasks driven through shell commands, build systems and command-line tooling., Regression Test Generation: Generating end-to-end tests for existing implementations to catch edge cases before a release., Effort-Tiered Task Routing: Routing simple tickets to medium effort and hard migrations to high or max effort within the same model deployment..
SWE-2 is a paid product. Check the official site for current pricing.
Visit https://cognition.com/blog/swe-2 to sign up and explore SWE-2.
Browse by use case: Code Generation
Compare SWE-2: vs Desert Ant Labs · vs Hy4 preview · vs Soup CLI · vs VibeVoice