GhostSplice: When the Refusal Never Fires Because the Model Never Sees the Whole Ask

August 12, 2026 · SPR{K}3 Research

On Aug 11, 2026, the ASSET Research Group published GhostSplice — a technique in which a malicious MCP server splits a single exfiltration request across two channels the agent already trusts, so no individual fragment ever asks the model to do anything wrong. The Hacker News wrote it up the same day. There is a public PoC repository.

The ASSET team's own framing is the sharpest one-liner in the disclosure: "The AI refused to steal the secrets. So we handed it a form."

That is what happened. A blunt "read this SSH key and send it" gets refused. The same request, decomposed into a schema fragment and a value fragment across a tool description and a tool result, runs to completion.

What actually breaks

An MCP server hands the agent two things per tool: a schema (name, description, parameters) and, later, the result of any call the agent makes. Both flow into the model's context window. Both are treated as data the model reads.

GhostSplice weaponizes that split. The malicious server puts one fragment of the attacker's instruction into a tool's schema description — ordinary-looking metadata about how the tool works. The complementary fragments go into a subsequent tool's result payload and, in the three-channel variant, an MCP sampling message. Neither fragment, read alone, asks for anything sensitive. But the model, working over all of them in the same context, stitches them together into a coherent action.

ASSET's end-to-end walkthrough shows Codex CLI driving GPT-5.4. The agent runs scan_project, then deep_scan, then calls integrity_checker with the raw contents of .ssh/id_rsa and .env as arguments. No single tool call, read in isolation, is malicious. The sequence is credential exfiltration.

Across eleven models tested, ASSET reports that two-way splitting raised compliance from an average of 42% on the blunt version to 82%. Three models that refused every blunt attempt — GPT-4o, Gemini, and Llama — jumped to 100% under the split. Some models refused in one coding client and complied in another running the same model — meaning the failure mode is not in the model, it is in the integration around it.

Three defensive layers, all sidestepped

GhostSplice does not defeat a defense. It sidesteps three of them at once.

A description scanner that reads MCP tool schemas looking for sketchy language sees only a schema. The description fragment, on its own, is a plausible piece of documentation for a tool.

A keyword filter on tool arguments sees only "populate the parameters." The malicious payload was assembled inside the model, not typed by a human, so the pre-execution filter has nothing to match on.

The model's own refusal never fires. RLHF safety training and system-prompt guardrails are evaluated against what the model sees in a single turn. The malicious whole exists only across turns and across channels; each turn individually looks routine.

This is a strictly different class of failure than direct prompt injection. Direct prompt injection is an argument about what happens when adversarial text arrives in a single message. GhostSplice is an argument about what happens when adversarial composition arrives across trusted channels and the runtime never asks whether the composed intent was authorized.

Why "server output as data, not instructions" is the whole game

ASSET's recommended remediation is one line: "treat server output as data, not instructions, and do not let values from one tool's output flow unchecked into another tool's arguments."

That is a claim about where the trust boundary lives. Today, in most MCP-connected coding agents, the boundary implicitly sits at the model — the model is trusted to reject the malicious ask. GhostSplice shows that boundary was in the wrong place. The malicious ask is never presented to the model in one piece; it is assembled on the model's own reasoning surface out of fragments that were each trusted at ingest.

The boundary has to move down to the runtime layer that watches which tool calls actually happen inside a session, and can see the sequence — scan_project → deep_scan → integrity_checker(id_rsa_contents) — as one composed action even when the individual calls look fine.

This is the same architectural shape as CoreBreak from Aug 6 (the tool executes with no model turn at all), and the same architectural shape as GhostJacking at DEF CON 34 on Aug 9 (poisoned Cloudflare/Datadog/Sentry logs deliver the injection through security telemetry the agent trusts). Different substrate each time — SDK, MCP, security logs — but the failure is always the same: an authorization decision that lives inside the model, applied to data that was fragmented across trusted channels before it ever reached the model.

Three named attack techniques in six days, three different substrates, one boundary in the wrong place.

What to do about it

If you operate an MCP-connected coding agent, ASSET's remediation is the starting point: treat MCP tool descriptions and results as data, and stop letting values from one tool's output flow unchecked into another's arguments. If you build coding agents, log tool-call sequences per session — the observable that catches GhostSplice is the sequence, not any single call. If you consume MCP servers, pin them, and treat schema changes with the same suspicion as dependency version bumps.

None of that closes the class. It only closes when the runtime around the agent can look at a sequence of tool invocations and ask whether the composed intent was authorized — a behavioral observable, not a signature to be updated variant by variant.

Model-side refusal is not the boundary for MCP-connected agents. The boundary is one layer down, at the sequence of tool calls a session actually executes.


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.