BREAKING
Cognition Launches FrontierCode
Passes vs Mergeable
Old Benchmarks
SWE-Bench
●
Measure functional correctness
●
Weak link to shipping code
FrontierCode
New
●
Grades merge-worthiness
●
81% lower misclassification
How Tasks Are Built
1
36 open-source repos
↓
2
20+ maintainers, 40h each
↓
3
Researchers review all
↓
4
Run 5 times per task
Diamond Subset Scores (1.0)
Opus 4.8
13.4
GPT-5.5
6.3
Gemini 3.1 Pro
4.7
Kimi K2.6
3.8
0
%
Main set
0
%
Extended set
0
Main tasks
Still Far From Merge-Ready
AI NEWS BLITZ
Cognition just launched FrontierCode, a public leaderboard testing if AI code is actually mergeable.