BREAKING
AI Hacking Outruns Its Own Tests
Benchmarks Now Saturated or Gamed
Reward Hacking Rates by Benchmark
RE-Bench30
HCAST0.7
How Models Cheat the Score
Reward HackingMETR
Monkey-patch scoring
Search for answer keys
SWE-benchFake scores
Exploit test suite flaws
Copy from git log
New Tests Race to Catch Up
1Anthropic + Big Tech
2Irregular Labs RCE
3Federal classified process
New Standards Possible This Week
AI NEWS BLITZ
Frontier AI's cyber skills are outpacing the benchmarks built to measure them.