On July 8, 2026, OpenAI retracted its earlier endorsement of SWE-Bench Pro, a widely used benchmark for measuring AI coding ability, saying roughly 30% of its tasks are broken and that it "no longer reliably measures frontier coding capability."
July 8, 2026 · OpenAI
OpenAI Pulls Its Backing of a Top Coding Benchmark
An audit found roughly a third of SWE-Bench Pro's tasks are broken. OpenAI says it "no longer reliably measures frontier coding capability" — the second benchmark it has abandoned in five months.
34.1%
of tasks judged broken by human reviewers (249 of 731)
27.4%
flagged broken by automated analysis (200 tasks)
~70%
"noise ceiling" caused by broken tasks
Scores Blew Past the Signal in 8 Months
SWE-Bench Pro top-model results climbed from 23.3% to 80.3% — pushing far above the ~70% ceiling where broken tasks make numbers meaningless.
- - - ~70% noise ceiling: everything above is unreliable
Repeated Benchmark Turnover
SWE-bench Verified
Discontinued Feb 2026
→
SWE-Bench Pro
Adopted Feb 2026
→
Endorsement withdrawn
Jul 2026
Harder, more realistic
Pro's commercial repositories stayed stringent — GPT-5 scored just 14.9% and Claude Opus 4.1 17.8% , reflecting genuine difficulty.
Why tasks break
Overly strict tests, underspecified or misleading prompts, and low test coverage — correct fixes still fail hidden tests.
Because benchmarks feed OpenAI's Preparedness Framework for safety and capability decisions, measurement reliability matters beyond ranking. Saturation and contamination remain an open challenge for evaluating AI coding agents.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…