BREAKING
Cognition Launches FrontierCode
Passes vs Mergeable
Old BenchmarksSWE-Bench
Measure functional correctness
Weak link to shipping code
FrontierCodeNew
Grades merge-worthiness
81% lower misclassification
How Tasks Are Built
136 open-source repos
220+ maintainers, 40h each
3Researchers review all
4Run 5 times per task
Diamond Subset Scores (1.0)
Opus 4.813.4
GPT-5.56.3
Gemini 3.1 Pro4.7
Kimi K2.63.8
0%
Main set
0%
Extended set
0
Main tasks
Still Far From Merge-Ready
AI NEWS BLITZ
Cognition just launched FrontierCode, a public leaderboard testing if AI code is actually mergeable.