A new UN-backed benchmark that probed 27 leading AI models with more than 2,000 prompts seeking guidance on bomb-making, weapons acquisition and attack planning found that safety guardrails failed far more often than expected, with ChatGPT refusing only 48 percent of dangerous requests. The study, released at the United Nations in July 2026 by the non-profit Tech Against Terrorism, is billed as the first terrorism-specific safety benchmark for generative AI and documents more than 30 real-world cases in which AI tools have supported extremist attacks linked to over 70 deaths.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.