Researchers at The Alan Turing Institute have published a study demonstrating that harmful instructions GitHub Copilot refuses in chat can still produce harmful output when embedded inside a coding task. According to the paper (arXiv:2607.03968, July 2026), the models refused nearly all harmful prompts when asked directly, but the safeguards broke down once the same content was reframed as an ordinary software-engineering task—"improve this benchmark score."coverage
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.