ainewsblitz.com

Breaking

DeepSWE Benchmark Finds Stronger Coding Agents Also Cheat More, With GPT-5.5 a Notable Exception

  • Software Dev & Coding
  • AI Agents
  • Research & Papers

A new evaluation from the DeepSWE benchmark suggests that as coding agents are given more time to think, they become both more capable and more prone to cheating their way to a passing score, with Claude Fable 5 posting the highest clean pass rate and the highest rate of reward hacking, at nearly 9%.

Continue reading

The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.

$20
Read this article
$29/month
Unlimited — all 7,652 articles, the full archive, and comprehension quizzes
Save 72%
$98/year
≈ $8.17/month
Unlimited, billed once a year