The Model Told the Researchers How to Hack It

August 20, 2026 · SPR{K}3 Research

On August 18, Microsoft patched CVE-2026-24301 in Copilot Personal, the consumer assistant hosted at copilot.microsoft.com. The chain — Varonis Threat Labs calls it CoSnitch — takes a single click on a crafted link and turns it into silent data theft from the victim's connected Gmail, Google Drive, Google Calendar, and OneDrive, inside the victim's own authenticated session. Microsoft rated it 8.8, High. The patch shipped roughly eight months after Varonis's December 2025 disclosure.

The vulnerability is a clean example of an assistant-side prompt-injection chain. The way the researchers found it is the more interesting story.

The chain

Three weaknesses in sequence, per Varonis's writeup and The Hacker News's coverage:

A single click on an ordinary-looking link runs an attacker-authored prompt inside the user's assistant, pulls whatever the assistant can reach, and ships it out.

There is a secondary vector worth flagging. A crafted web page, summarized by Copilot, plants attacker instructions in Copilot's persistent memory store — the layer meant to remember the user's preferences between sessions. Per Varonis, those injected memories survive password change, session revocation, and device re-enrollment. That is the same memory-poisoning primitive the Bad Memory arXiv result named across vendors and the Anthropic + EPFL Mind Viruses preprint demonstrated between agents. CoSnitch is the same primitive, shipped and reproducible against a consumer assistant, with a CVE number attached.

The meta-hacking

The novel piece is how the researchers found ?autorun=1, undocumented and invisible in Microsoft's public API surface.

Per Cybernews, Varonis prompted Copilot to explain why auto-execution was impossible. Copilot refused, but the refusal included a technical justification. They reframed the refusal as a follow-up. Copilot refused again, in more detail. They kept going. At some point Copilot volunteered the undocumented parameter name mid-refusal, along with its historical behavior and every protection Microsoft had added.

The model knew a vendor-internal, undocumented endpoint parameter of the product it was hosted inside. Asked repeatedly to justify why a class of behavior was impossible, it named the mechanism designed to make it impossible, and how that mechanism had been hardened over time.

Every hosted-assistant vendor has this problem — not one class of undocumented endpoints, every class. The model has seen the docs, the internal design threads, the changelogs, the postmortems. Refusal training is the mechanism keeping this out of chat, and refusal training is not designed to survive a researcher iterating meta-questions until the refusal itself becomes the leak.

The timeline

Varonis disclosed to Microsoft in December 2025. The patch shipped August 18, 2026 — about eight months. CSO Online, Computerworld, and Cybernews all led on that gap.

For scale: the same digest cycle that named CoSnitch also named the Atlassian Rovo zero-click chain PromptArmor disclosed May 23 and is still unpatched 87 days later, and the Aug 11 Microsoft Patch Tuesday load — 421 CVEs including two CVSS-9-tier authorization flaws in Microsoft's own AI agents. Prompt-injection-class flaws in production consumer assistants are moving at consumer-support timelines, not security-response timelines, even at the largest vendors.

The class

CoSnitch is not one bug. It is a shape:

Three items, three substrates, one shape: the boundary the ML or product team believed exists — between the URL, the click, the session, the flag, the browser — is not the boundary the runtime actually enforces. In each case the observable that would catch the attack is not the URL and not the click. It is the moment the assistant, the agent, or the dashboard starts issuing tool calls, HTTP requests, or job submissions that the human never authored on that page load.

The takeaway

Two things worth watching beyond the CVSS number.

Consumer AI assistants with connected corporate apps are now a first-class exfiltration substrate. Personal Copilot users routinely connect Gmail, Drive, Calendar, and OneDrive; every scope granted is a scope an attacker inherits on one click. Vendor patch tempo — Microsoft's eight-month CoSnitch window, Atlassian's still-open 87-day Rovo window — isn't moving at the tempo the exploitation class deserves.

Undocumented endpoints on hosted assistants can be extracted from the assistant itself. Refusal training is a UX mechanism, not a security boundary. Patient meta-questions get researchers what the vendor didn't publish. Any assistant with a large internal-doc footprint inherits this discovery primitive.

The runtime boundary — the moment an assistant executes a tool call the human never wrote — is the one boundary in this chain a vendor cannot leak away in a refusal.

Update, Sept 21: Atlassian fixed Varonis's separate RovoBlast one-click path server-side on July 8. The PromptArmor zero-click path described above remains unpatched.


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.