AI Briefing
KO

Cognition Releases SWE-2 Based on Kimi K3

·2026.09.12 10:30

Key point

Cognition released SWE-2, based on Kimi K3, achieving less than a 1 percentage point difference from Fable 5.1 on the FrontierCode benchmark while reducing costs by 64%.

1 / 10

Details

Cognition released SWE-2, a coding model based on Moonshot AI's open-weight model Kimi K3 (2.8T parameter MoE). This model maximizes performance through post-training and reinforcement learning (RL), and is immediately available in Devin Desktop and CLI.

Performance and Cost Efficiency

On the FrontierCode 1.1 Main benchmark, SWE-2 scored 50.0%, showing a difference of less than 1 percentage point from Claude Fable 5.1 (50.9%). Meanwhile, the average cost per task was approximately $1.2, representing a 64% reduction compared to Fable 5.1 medium ($3.28). On Terminal-Bench 2.1, it recorded the highest score among the models in the original comparison table at 92.8%, but scored 27.3% on Terminal-Bench 4, lower than frontier models, indicating performance differences depending on the complexity of agentic tasks.

Training Methodology and Optimization

Cognition adopted an approach that simultaneously trains all Reasoning Effort levels (medium/high/max) in a single RL run. For Pareto frontier optimization, they designed a reward function considering success rate (S) and cost (C) (R = S - λ_e C), setting the cost penalty coefficient (λ_e) according to the local slope of the frontier curve. In a comparison running 100 tasks from FrontierCode 1.1 Main three times per model, SWE-2 medium reduced the average number of steps from 127 to 53 (a 58% reduction) compared to SWE-1.7, and the average cost was 81% lower.

Infrastructure and Deployment

To serve RL for the 2.8 trillion parameter scale, they introduced Prefill Delayer and Speculative Decoding (DSpark) to improve throughput. NVFP4/FP8 kernels and QAT (Quantization-Aware Training) were applied to enhance inference efficiency. SWE-2 is deployed exclusively for the Devin product line, and the weights are not released.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.