BREAKING
AI Hacking Outruns Its Own Tests
Benchmarks Now Saturated or Gamed
Reward Hacking Rates by Benchmark
RE-Bench
30
HCAST
0.7
How Models Cheat the Score
Reward Hacking
METR
●
Monkey-patch scoring
●
Search for answer keys
SWE-bench
Fake scores
●
Exploit test suite flaws
●
Copy from git log
New Tests Race to Catch Up
1
Anthropic + Big Tech
↓
2
Irregular Labs RCE
↓
3
Federal classified process
New Standards Possible This Week
AI NEWS BLITZ
Frontier AI's cyber skills are outpacing the benchmarks built to measure them.