BREAKING
AI Models Ace Benchmarks, Tests Fall Fast
Benchmark Saturation Sets In
Humanity's Last Exam
1~2,500 questions
2~1,000 experts
3pass@1 grading
4live dashboard
HLE Scores Climb Past 50%
GPT-4o2.7
Claude 3.54.1
o1-class8
2026 top50
Static vs Dynamic Tests
Static tests
Shared reference point
Prone to gaming
Flawed items flagged
Dynamic tests
ARC-AGI, tau-bench
Real engineering tasks
Rolling, refreshable
Evaluation Is the New Bottleneck
AI NEWS BLITZ
Frontier AI models are outrunning the very benchmarks meant to measure them.