BREAKING
Anthropic Details Fable 5 Cyber Safeguards
Classifiers Sort Requests Into 4 Tiers
1
Prohibited use
↓
2
High-risk dual use
↓
3
Low-risk dual use
↓
4
Benign use
0
%
Blocks reported bypass
0
%
Sessions triggered
Jailbreak Severity Framework: 4 Criteria
Capability
●
Capability gain
●
Breadth of gain
Exploitability
●
Ease of weaponization
●
Discoverability
0
$
Per 1M input
0
$
Per 1M output
0
M
Context tokens
A Common Language for Cyber Safety
AI NEWS BLITZ
Anthropic has published its Fable 5 cyber safeguards and a draft jailbreak grading framework.