ainewsblitz.com

Breaking

Cognition Launches FrontierCode Leaderboard to Measure Whether AI-Written Code Is Actually Mergeable

  • Software Dev & Coding
  • Research & Papers
  • Foundation Models

Cognition, the company behind the Devin coding agent, has launched a public leaderboard for FrontierCode, a benchmark that scores AI models not on whether their code passes tests but on whether a human maintainer would actually merge it into production. The live page publishes full scores — including newly added results for Grok 4.5 and Inkling — alongside the complete methodology and sample tasks.

Continue reading

The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.

$20
Read this article
$29/month
Unlimited — all 6,426 articles, the full archive, and comprehension quizzes
Save 72%
$98/year
≈ $8.17/month
Unlimited, billed once a year