On July 8, 2026, Anthropic, in a joint study with AE Studio, released GRAM (Gradient-Routed Auxiliary Modules), a training method that isolates "dual-use" knowledge such as virology and cybersecurity into removable modules within an AI model. By separating potentially dangerous knowledge during training and deleting the module at inference time, the approach aims to create what amounts to an "off switch" for such capabilities. Research overview
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.