BREAKING
OpenAI Drops SWE-Bench Pro Endorsement
0
%
auto-flagged (200 tasks)
0
%
human-flagged (249 tasks)
0
public tasks
Scores climbed to a noise ceiling
Launch GPT-5
23.3
8 months later
80.3
Noise ceiling
70
Repeated benchmark turnover
1
Verified dropped Feb 2026
↓
2
Switched to Pro
↓
3
Pro endorsement pulled
Why the tasks break down
Test flaws
●
Overly strict tests
●
Low test coverage
Prompt flaws
●
Underspecified prompts
●
Misleading prompts
Saturation remains an open challenge
AI NEWS BLITZ
OpenAI has pulled its backing of the SWE-Bench Pro coding benchmark.