BREAKING
OpenAI: 30% of SWE-Bench Pro Broken
0
tasks
0
repositories
0
%
broken tasks
How Tasks Break Down
1
Hidden requirements
↓
2
Contradictory instructions
↓
3
Overly strict tests
↓
4
Incomplete grading
Two Benchmarks, Same Fate
SWE-bench Verified
Dropped Feb 2026
●
500 human-verified tasks
●
Dropped over contamination
SWE-Bench Pro
Dropped Jul 2026
●
1,865 tasks by Scale AI
●
Dropped: 30% broken
The Benchmark Cycle Repeats
Reliable Evaluation Stays Hard
AI NEWS BLITZ
OpenAI says nearly a third of SWE-Bench Pro tasks are broken and drops its endorsement.