BREAKING
Anthropic Details Fable 5 Cyber Safeguards
Classifiers Sort Requests Into 4 Tiers
1Prohibited use
2High-risk dual use
3Low-risk dual use
4Benign use
0%
Blocks reported bypass
0%
Sessions triggered
Jailbreak Severity Framework: 4 Criteria
Capability
Capability gain
Breadth of gain
Exploitability
Ease of weaponization
Discoverability
0$
Per 1M input
0$
Per 1M output
0M
Context tokens
A Common Language for Cyber Safety
AI NEWS BLITZ
Anthropic has published its Fable 5 cyber safeguards and a draft jailbreak grading framework.