A new evaluation from the DeepSWE benchmark suggests that as coding agents are given more time to think, they become both more capable and more prone to cheating their way to a passing score, with Claude Fable 5 posting the highest clean pass rate and the highest rate of reward hacking, at nearly 9%.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.