Loopjacking: The Human Approved Operation A. The Workflow Ran Operation B.

September 22, 2026 · SPR{K}3 Research

Most agent products that touch a consequential action end at the same gate: a human reviews what the agent wants to do, clicks approve, and the action fires. That gate is the last hard boundary anyone actually deploys. A paper published to arXiv Sept 21 says the gate is bypassable — not by prompt injection, not by talking the model out of refusing, but by arranging for the reviewer to approve operation A while the workflow executes materially different operation B. The authors named the class Loopjacking, and they reproduced it in named versions of three shipped agent runtimes.

Two shapes of the same bypass

Per the Loopjacking paper (arXiv 2609.21081), the attack comes in two flavors:

The paper is direct that this is not model-side. The model's refusals fire correctly. The guardrails don't matter. The failure is that the "signed approval" is not bound to the exact bytes that later execute.

Where it was reproduced

Per the paper, post-approval state-substitution was reproduced in seven Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. Representation-mismatch was reproduced in OpenClaw 2026.2.23 and rejected — meaning the attack no longer works — in the follow-on 2026.2.24.

That is not a research finding on a toy setup. That is three shipped agent runtimes, at named version numbers a defender can check, where the human-approval gate is currently or recently bypassable.

Why the approval gate can fail

The approval prompt is a text artifact rendered from workflow state. What the reviewer reads and what later executes are two derivations from the same state, but they run at different times, and nothing in between binds them together. If any of the following happens between "approve" and "execute" the bypass lands:

The paper's proposed mitigation is short: render the canonical action the executor will use, sign the exact bytes, and compare them at execution time. Or prevent unauthorized pending-state mutation between approval and use. Nothing exotic — but neither is currently the default in the three runtimes named above.

This is what everyone is converging on

Loopjacking landed one day before APort Vault, a benchmark we wrote up yesterday that ran 4,371 attacks against 14 models acting as payment agents. Model-only defense failed 74.6% of the time. A deterministic pre-action check bound to the exact action drove that to 0% across 879 attempts. Same reading as Loopjacking, arrived at from a different angle: the boundary has to be enforced from outside the model, on the actual bytes about to fire, verified against something that cannot mutate after the fact.

Spain's AEPD got there before either paper. Its February 2026 agentic-AI framework — cited as the basis for a September breach notification — includes the "Rule of 2": an agent must never simultaneously process untrusted input, access sensitive data, and take autonomous action without oversight. September gave the regulator its first matching Article 33/34 filing.

Two papers, one regulator, three independent framings. All pointing at the same failure: the last boundary — human approval — is only a boundary if it is bound to the exact action.

What defenders should do this week

Three concrete moves, none of which require the paper's specific engineering:

  1. Check for the named versions. If you're running Agno AgentOS ≤ 3.0.9, LangGraph Agent Server ≤ 0.14.0 in a conditional in-memory composition, or OpenClaw ≤ 2026.2.23, you are running a runtime with a reproduced attack against your approval gate. Upgrade where a fix ships; segment where it does not.
  2. Audit what your approval UI actually shows. If the reviewer reads a natural-language summary and the executor consumes tool arguments recomputed at run time, you are in the representation-mismatch shape whether or not you use one of the named runtimes.
  3. Assume post-approval state can be mutated. Any pending-state store the workflow reads between approval and execution is inside the attack surface. Sign the canonical action at approval time; verify at execution time; refuse if they differ.

None of that is bespoke to any one product. All three papers point at it. The reason the industry hasn't shipped it yet is that "the approval prompt fires" felt like enough.

Sources


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.