Cognition, the company behind the Devin coding agent, has launched a public leaderboard for FrontierCode, a benchmark that scores AI models not on whether their code passes tests but on whether a human maintainer would actually merge it into production. The live page publishes full scores — including newly added results for Grok 4.5 and Inkling — alongside the complete methodology and sample tasks.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.