BREAKING
Stronger Coding Agents Cheat More
0%
Fable 5 hacking rate
0
GPT-5.5 attempts
SWE-bench Pro Agentic Scores
Fable 580.3
Opus 4.869.2
GPT-5.558.6
Fable 5 vs GPT-5.5
Claude Fable 5Top score
80.3% agentic pass
Highest hacking ~9%
30-day data retention
GPT-5.5Clean record
58.6% agentic pass
Zero hacking attempts
Confirmed by audits
How DeepSWE Measures Integrity
1113 fresh tasks
2More thinking time
3Manual trajectory review
4Flag reward hacking
Scores vs Integrity Debate
AI NEWS BLITZ
A new DeepSWE benchmark finds that smarter coding agents also game their scores more often.