Drawing on the Future of Life Institute's (FLI) "AI Safety Index: Winter 2025," a report says leading AI firms including Anthropic, OpenAI, Google DeepMind and Meta are weakening or withdrawing key safety commitments as their models grow more capable.
December 2025 · Future of Life Institute
AI Safety Index: Eight Labs, One Failing Grade
An expert panel scored eight leading AI companies across six domains and 35 indicators. The top mark was a C+. None scored above 2.67 out of 4.3 — and several are quietly walking back the safety pledges they once made.
C+
Highest grade awarded (Anthropic, 2.67 / 4.3)
8
Companies evaluated across 6 domains, 35 indicators
0
Firms with a clear plan to control superintelligence
The Scorecard
Score out of 4.3 — column heights drawn to scale
"Moving the Goalposts"
Anthropic, OpenAI, Google DeepMind and Meta have softened or deleted earlier pledges. The weakest domain across all firms was Existential Safety — measures against loss of control from superintelligence.
PAUSE PLEDGES
Commitments to halt at dangerous thresholds were watered down or removed.
MILITARY USE
Once prohibited — now Anthropic, OpenAI and Google work with the military.
TEST BEHAVIOR
Models showed blackmail and deception in controlled test settings.
⚠ THE WARNING
Voluntary frameworks — the 2023 White House pledge and 2024 AI Seoul Summit "kill switch" commitments — are eroding before permanent regulation is in place.
◇ THE PUSHBACK
Companies argue that with open-weight models, the deploying firm decides on fine-tuning and safety controls — not the original developer.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…