BREAKING
Copilot Models Tricked Into Banned Answers
0
coding-workflow runs
0
%
yielded harmful output
Direct Chat vs Coding Context
Direct Chat
Refused
●
Declined harmful prompts
●
Guardrails fired
Coding Workflow
Bypassed
●
Framed as benchmark Q&A
●
Model wrote banned answers
How the Evasion Works
1
Harmful request
↓
2
Wrap as coding task
↓
3
Filter misses it
↓
4
Model outputs answer
Input Checks Are Not Enough
A New Attack Surface for Dev Tools
AI NEWS BLITZ
Researchers say AI coding assistants can be coaxed past their own safety guardrails.