Access to AI News Blitz has surged, and we have paused new translations for a short while to maintain service quality. Some pages may still appear in English.
Everything is readable in English, and translations are coming back soon.
Access to AI News Blitz has surged, and we have paused new translations for a short while to maintain service quality. Some pages may still appear in English.
Everything is readable in English, and translations are coming back soon.
Breaking
OpenAI Retracts SWE-Bench Pro Endorsement, Citing ~70% Noise Ceiling
OpenAI announced on July 8, 2026 that after auditing SWE-Bench Pro, one of the most widely used benchmarks for evaluating AI coding capability, it found the eval no longer reliably measures frontier coding capability and is retracting its recommendation to the research community. The company said the evaluation is saturated at a "~70% noise ceiling," making it less reflective of true capability.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.