The Guardrail Reads the Wire. The Model Reads the Cleartext.
Adversa's "Cryptographic Context Injection" gets Grok to exfiltrate user data — and Gemini to answer questions its safety filter would normally refuse
On Aug 20-21, Adversa AI researcher Rony Utevsky went public with an indirect prompt-injection technique he calls Cryptographic Context Injection (CCI). Per The Register and SecurityWeek, it works against xAI's Grok 4.5 Fast on grok.com and Google's Gemini. Reported to xAI Jun 3, 2026, coordinated-disclosure attempts Aug 4 and Aug 10, published Aug 20-21 with no fix confirmation. About 78 days end to end. Reproduction: 20 attempts since June, ~40% success rate; the coverage demo reproduced once on Aug 19.
The Grok PoC ends with the user's name, approximate location, subscription tier, and chat history landing in an attacker webhook — no click, no confirmation, no visible warning. Against Gemini, the same technique produced content the filter normally blocks — Adversa cited weapons-adjacent instructions.
The technique in one paragraph
The attacker's page carries three things: an AES-256-GCM ciphertext of the malicious instructions, a plaintext description telling the assistant the page is encrypted and how to decrypt it, and the key material (or a PBKDF2 recipe to derive it). The user asks Grok (or Gemini) to summarize the page. The assistant hands the ciphertext to its own code-execution sandbox to run PBKDF2 + AES-256-GCM against the supplied key. The decrypted plaintext comes back into the assistant's context as trusted output from its own tool — no longer an untrusted page fragment. From there the assistant executes what the decoded content says: a URL fetch that packs the user's identifiers into the request path and delivers them to an attacker webhook. This distinguishes CCI from earlier work like CipherChat and CodeChameleon, which used substitution / XOR / base64 — schemes the model decodes natively without an interpreter. AES-256-GCM cannot be decoded that way; a content classifier would need to run real cryptography at inspection time to see the payload, and none do.
Why the guardrail didn't catch it
The guardrail runs on the wire. It sees the plaintext description and the AES-256-GCM blob. Content filters pattern-match on text, and the blob doesn't match anything harmful because it isn't readable yet. By the time the payload is plaintext, it has passed through the assistant's own interpreter and re-entered context as trusted tool output. There is no second inspection stage. The gap between "what the filter reads" and "what the assistant acts on" is the attack surface.
Same shape as Microsoft's Copilot CoSnitch chain: guardrails inspected the URL, the model reasoned over the auto-executed prompt inside its session. Same shape as PromptArmor's Atlassian Rovo disclosure: admin-facing filters gated one tool, the assistant reached the data through another.
The capability that makes this land
The delivery vehicle is a general capability: a code-execution sandbox that runs whatever Python the assistant asks it to. Every frontier assistant that ships a code interpreter has it. That makes the usual patch approach — "detect and refuse encoded payloads" — hard. You cannot refuse ciphertext without breaking legitimate use (papers with cipher tables, historical cryptography, captured-traffic writeups). Capability improvements make CCI stronger, not weaker: a more capable interpreter handles more cipher schemes, more key derivations, more container formats. The scaling trend is on the attacker's side.
The exfiltration path
The URL-fetch tool is not a bug. Grok opens URLs, Gemini looks things up. The attacker borrows the assistant's existing URL-fetch and points it at their server, payload in the path. Adversa's four fields — name, approximate location, subscription tier, current chat history — are already in session state; exfil is just moving them through the assistant's own tool. Same shape as CoSnitch and Rovo.
Cross-vendor
Adversa got the same technique past Gemini's safety filter — output the filter would normally block. Two vendors, two independent guardrail stacks, one technique. That is a class vulnerability against how frontier-model safety filters are architected, not one vendor's oversight.
Where this leaves defenders
Wire guardrails still catch a lot but cannot be the floor for prompt-injection defense: an assistant that can hand its interpreter a decryption task moves the payload past them. The observable that catches CCI is downstream of both filter and decryption step — the moment the reasoning trace produces a tool call the human never authored, a URL fetch to an unfamiliar host with session identifiers in the path. A runtime boundary catches that without inspecting content.
Concretely, for an assistant that summarizes web content and calls URLs:
- Content-filter passes are necessary but not sufficient — the filter cannot see what the interpreter will produce.
- Constrain what URL-fetch can reach. Prefer an outbound-host allowlist. A fetch to a host the assistant has never touched, mid-summary of an untrusted page, is the observable.
- Log the transition from "user asked to summarize" to "assistant issued a request the user didn't ask for."
- Vendor patch tempo is weeks-to-months. Rovo (PromptArmor zero-click): 87 days unpatched at disclosure, still open. CoSnitch: ~8 months from Varonis report to Microsoft patch. CCI: reported Jun 3, ~78 days without a fix at Adversa's Aug 21 disclosure, still unpatched at Grok in most recent reporting.
The takeaway
CCI rides on the interpreter the assistant is supposed to have. Guardrail sees ciphertext, interpreter returns cleartext into trusted context, user sees the summary, attacker sees the URL fetch land at their webhook.
The boundary that stays open when the content filter can't see the payload is the one three weeks of adjacent disclosures — CoSnitch, Rovo, CCI — all point at: watch what the assistant's tools do the moment the reasoning trace takes a hand-off from something the user did not write.
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.