On July 8, 2026, OpenAI announced that roughly 30% of the tasks in SWE-Bench Pro, one of the most widely used benchmarks for evaluating AI coding capability, are broken and that the benchmark can no longer reliably measure frontier coding performance. The company retracted its earlier recommendation of the benchmark as a leading coding evaluation for the research community.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.