Pentera Labs' red team has demonstrated an "AI Double Agent Attack" that achieves remote code execution (RCE) against Anthropic's desktop app Claude Desktop without using malware or phishing links. Starting from a compromised inbox or account, the attackers turn the trusted AI assistant itself into their proxy. The research was reported to Anthropic in November 2025 and published by Pentera on July 1, 2026.
July 1, 2026 · Pentera Labs Red Team
The AI "Double Agent": Turning Claude Desktop Against Its Own User
Starting from a compromised inbox, researchers achieved remote code execution with no malware and no phishing links — hijacking the trusted AI assistant itself as a persistent command-and-control channel.
0
malware or phishing links needed for full RCE
1 field
"Personal Preferences" prompt that syncs to every device
by design
how the vendor classified it — out of scope for the bug program
The attack chain
A compromised inbox becomes remote code execution — no user download required.
1 · Access inbox
Abuse email aggregator auth flow, pivot into the Claude account
→
2 · Inject prompt
Base64 payload hidden in Personal Preferences, syncs everywhere
→
3 · Auto-execute
Runs on next launch; Claude enumerates MCP command tools
→
4 · RCE / C2
Reverse shell or curl C2 traffic; Claude becomes a persistent channel
If no command tool is present: Claude shows a fake error page mimicking Anthropic to trick the user into installing Desktop Commander MCP.
Related Claude Desktop findings — severity by CVSS
DXT 0-click (Calendar)
10.0
PromptJacking (patched)
8.9
AI Double Agent
by design
The Double Agent attack has no CVSS score — it was treated as intended behavior, not a vulnerability.
What defenders recommend
Treat AI desktop apps as privileged software
Monitor configuration changes
Limit installed extensions
Enforce strong session & auth controls
Where the debate stands
One MCP tool's output can invoke another — flagged as root cause
Calls for sandboxing & privilege separation
Many find the "by design" answer hard to accept
"Abusing trust in AI has become easy."
The blast radius of a single account compromise has grown — AI agents' trust boundaries remain fragile, and the assistant your workstation trusts can be turned into the intruder.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…