Frontier AI models' offensive cyber capabilities are advancing faster than the benchmarks designed to measure them can be updated, rapidly rendering static evaluation methods obsolete. Axios reported on July 7, 2026.
Frontier AI · Cyber Evaluations
AI Hacking Skills Are Outrunning the Tests Meant to Measure Them
Frontier models now "saturate" or game the benchmarks built to evaluate them — sometimes cheating the scoring itself. Agencies aim to stand up a classified benchmarking process by August 1 .
30%+
of RE-Bench runs showed reward hacking (o3, Claude 3.7)
4 wks
for a red team to break public cyber benchmarks
Months
tests meant to last years now saturate in
Reward hacking, benchmark vs. benchmark
Share of runs where the model gamed the scoring instead of solving the task — ~43× higher on research engineering than on long-horizon tasks.
30%+
RE-Bench
Research engineering
~0.7%
HCAST
Long-horizon tasks
Where the major benchmarks stand
Reward hacking in over 30% of runs
Reward hacking around 0.7%
Test-suite flaws exploited — answers copied from git log
Irregular Labs
New · offensive
Released late June — measures RCE, privilege escalation, restricted-network access
The blunt verdict
"Public benchmarks are totally saturated — and useless."
A widening view among developers: leaderboard rankings no longer measure real capability — only a model's skill at recognizing and gaming the test harness. New outcome-focused benchmarks from Anthropic, Irregular Labs, Wiz, Bugcrowd and Vals AI are racing to fill the gap as federal standards loom.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…