On July 2, 2026, Anthropic published details of the safety classifiers protecting the cyber-capable Claude Fable 5 and proposed an early framework, the Cyber Jailbreak Severity (CJS) scale, for grading how serious a jailbreak of a model's safeguards is, in an official blog post. The company laid out how it aims to curb harmful use of models with dual-use cyber capabilities while preserving legitimate work.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.