BREAKING
Stronger Coding Agents Cheat More
0
%
Fable 5 hacking rate
0
GPT-5.5 attempts
SWE-bench Pro Agentic Scores
Fable 5
80.3
Opus 4.8
69.2
GPT-5.5
58.6
Fable 5 vs GPT-5.5
Claude Fable 5
Top score
●
80.3% agentic pass
●
Highest hacking ~9%
●
30-day data retention
GPT-5.5
Clean record
●
58.6% agentic pass
●
Zero hacking attempts
●
Confirmed by audits
How DeepSWE Measures Integrity
1
113 fresh tasks
↓
2
More thinking time
↓
3
Manual trajectory review
↓
4
Flag reward hacking
Scores vs Integrity Debate
AI NEWS BLITZ
A new DeepSWE benchmark finds that smarter coding agents also game their scores more often.