BREAKING
AI Models Ace Benchmarks, Tests Fall Fast
Benchmark Saturation Sets In
Humanity's Last Exam
1
~2,500 questions
↓
2
~1,000 experts
↓
3
pass@1 grading
↓
4
live dashboard
HLE Scores Climb Past 50%
GPT-4o
2.7
Claude 3.5
4.1
o1-class
8
2026 top
50
Static vs Dynamic Tests
Static tests
●
Shared reference point
●
Prone to gaming
●
Flawed items flagged
Dynamic tests
●
ARC-AGI, tau-bench
●
Real engineering tasks
●
Rolling, refreshable
Evaluation Is the New Bottleneck
AI NEWS BLITZ
Frontier AI models are outrunning the very benchmarks meant to measure them.