BREAKING
OpenAI Drops SWE-Bench Pro Endorsement
0%
auto-flagged (200 tasks)
0%
human-flagged (249 tasks)
0
public tasks
Scores climbed to a noise ceiling
Launch GPT-523.3
8 months later80.3
Noise ceiling70
Repeated benchmark turnover
1Verified dropped Feb 2026
2Switched to Pro
3Pro endorsement pulled
Why the tasks break down
Test flaws
Overly strict tests
Low test coverage
Prompt flaws
Underspecified prompts
Misleading prompts
Saturation remains an open challenge
AI NEWS BLITZ
OpenAI has pulled its backing of the SWE-Bench Pro coding benchmark.