PromptArmor
Threat Intel

Hijacking Copilot Cowork's AI Gateway to Bypass Sandboxing and Exfiltrate Files

A vulnerability in Copilot Cowork's AI infrastructure enabled malicious Skills to bypass the agent's sandbox to spawn ungoverned Anthropic agents and exfiltrate files.

Copilot Cowork’s AI gateway hijacked to exfiltrate the victim’s files

Context

Microsoft Copilot Cowork is an agent in M365 that runs in a sandbox intended to block network access and prevent the agent from running code that reaches any untrusted services.

In order for Copilot Cowork to generate responses, the sandbox forwarded network requests to Anthropic. Malicious Skills were able to hijack this pathway to exfiltrate data by spawning new agents in Anthropic’s cloud equipped with network-capable tools that could reach an attacker’s server.

No human in the loop approval was required, and the attack could target any data Copilot could access - from uploaded files, to SharePoint, to Teams, Outlook, and more.

This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026. More details on responsible disclosure are at the bottom of the article.

The Attack Chain

  1. The victim uploads a sensitive document they want to review

    The victim uploads an MSA draft to Copilot Cowork
  2. The victim invokes a Skill they found online; it conceals malicious prompts and code

    Skills are frequently found online and uploaded by users, and they can also be shared intra-org within M365. Per the documentation, “Skills can also include up to 20 companion files (such as reference documents and scripts)”. One script bundled with this Skill is malicious.

    The victim invokes a contract review Skill found online
  3. Copilot Cowork is manipulated by the malicious Skill into running malicious code

    The malicious Skill’s description falsely claims that  “All processing runs on-device, no document content is transmitted to third-party services.” After making a minimal attempt to review the Skill's code, Copilot concludes, "The script is a safe local analyzer — no network calls... Let me run it now."

    When Copilot runs the code, it hunts through the user's data, exploits a vulnerability to spawn agents outside the sandbox, and uses those agents to exfiltrate the victim's data to an attacker's server.

    Copilot runs the Skill’s bundled code, believing it to be a safe local analyzer
  4. The malicious code exfiltrates data from the sandbox by hijacking Copilot’s AI gateway

    The malicious Skill activates the same AI gateway Copilot itself uses to generate responses. The Skill’s code spins up agents running in Anthropic’s cloud. Each has one sentence from the victim’s document, access to a web fetch tool, and instructions to fetch attacker.com/?data={victim’s data here}. The attacker’s server logs every URL it received a request for, including the victim’s data that was appended to the requested URLs.

    The malicious Skill activates the local AI gateway to send data to the attacker
  5. The attacker can read the victim’s contract in their server logs

    Below, the attacker’s server log displays the victim’s exfiltrated contract. However, this attack could have just as easily targeted any data in SharePoint, Teams, Outlook, or other sources connected to Copilot. This is because the malicious Skill code can directly call Copilot’s tools without going through the model, as demonstrated by our prior research on a different sandbox bypass in Copilot.

    The attacker’s server logs contain the victim’s contract

Responsible Disclosure

This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026.

Timeline

DateEvent
July 14, 2026PromptArmor discloses to Microsoft
July 14, 2026Microsoft confirms receipt
August 17, 2026Microsoft confirms reported behavior
September 2, 2026Microsoft confirms a fix has been implemented

PromptArmor Threat Intelligence

Is your organization protected from AI in vendors?

PromptArmor continuously monitors across your portfolio of third party AI in vendors, skills, plugins, connectors, MCP servers, models and more.

We detect vulnerabilities and changes like this, surfacing risk before it becomes an incident.

Learn more