On July 8, 2026, Cognition released SWE-1.7, a software-engineering model for its autonomous AI engineer Devin, built on Moonshot AI's open-source Kimi K2.7 Code and enhanced through proprietary reinforcement learning (RL) for both capability and trustworthiness.
July 8, 2026 · Cognition
RL on top of RL: SWE-1.7 quadruples Devin's coding score
Cognition's new software-engineering model — built on Moonshot AI's open-source Kimi K2.7 Code and refined with proprietary reinforcement learning — jumps from 9.4% to 42.3% on its FrontierCode benchmark, closing in on US frontier systems at a fraction of the cost.
42.3%
FrontierCode 1.1 score — up from 9.4% on SWE-1.6
~1,000
Tokens/second on Cerebras infrastructure
$1.97
Cost per task on FrontierCode Main
FrontierCode 1.1 — score, drawn to scale
Column height ∝ benchmark %. SWE-1.7 lands between its base model and the frontier.
Full benchmark comparison (%)
Model
FrontierCode
Terminal-Bench
SWE-Multiling.
Kimi K2.7 Code
30.1
72.7
73.5
Trustworthiness: retraining the base model's behavior
Cognition's RL targets both capability and refusal behavior on sensitive tasks.
87%
of human-rights-sensitive tasks completed by base Kimi K2.7 Code — where other models refuse
SWE-1.7 refuses
After trustworthiness training, it behaves comparably to US frontier models on surveillance scenarios
The achievement
A 10+ point gain built on an already RL-trained baseline — evidence that RL can push past the post-training ceiling, delivered at ~1,000 tok/s and free to paying users through Aug 8.
The caution
Most figures are self-reported and await independent verification, with warnings of possible overstatement given contamination around SWE-Bench-style benchmarks.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…