BREAKING
'Context Bombs' Turn AI Attackers' Guardrails
How a context bomb works
1Plant string in canary
2Agent reaches decoy
3Guardrail triggers
4Attack aborts
0%
Baseline admin
0%
With bombs
Attack success drops sharply
Any path base91
Any path bombed15
Compromise base36
Compromise bombed1
Tailored triggers per model
Chinese modelstrigger
Politically sensitive prompts
In Chinese
Western modelstrigger
Biosecurity content
Blocks Opus, Gemini
Not a complete blockade
AI NEWS BLITZ
Researchers say they can stop AI hacking agents by triggering the models' own safety filters.