Loopjacking: The Human Approved Operation A. The Workflow Ran Operation B.
Most agent products that touch a consequential action end at the same gate: a human reviews what the agent wants to do, clicks approve, and the action fires. That gate is the last hard boundary anyone actually deploys. A paper published to arXiv Sept 21 says the gate is bypassable — not by prompt injection, not by talking the model out of refusing, but by arranging for the reviewer to approve operation A while the workflow executes materially different operation B. The authors named the class Loopjacking, and they reproduced it in named versions of three shipped agent runtimes.
Two shapes of the same bypass
Per the Loopjacking paper (arXiv 2609.21081), the attack comes in two flavors:
- Representation mismatch. Operation B is already encoded in the request at approval time, but the approval UI or serialized approval payload hides or misrepresents it. The reviewer sees A. The workflow later executes what's actually there.
- Post-approval state substitution. The reviewer sees A and approves A. Then, before the action fires, some mutable workflow state — pending task, tool arguments, cached plan — gets rewritten to B. The approval token still validates. The wrong action still runs.
The paper is direct that this is not model-side. The model's refusals fire correctly. The guardrails don't matter. The failure is that the "signed approval" is not bound to the exact bytes that later execute.
Where it was reproduced
Per the paper, post-approval state-substitution was reproduced in seven Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. Representation-mismatch was reproduced in OpenClaw 2026.2.23 and rejected — meaning the attack no longer works — in the follow-on 2026.2.24.
That is not a research finding on a toy setup. That is three shipped agent runtimes, at named version numbers a defender can check, where the human-approval gate is currently or recently bypassable.
Why the approval gate can fail
The approval prompt is a text artifact rendered from workflow state. What the reviewer reads and what later executes are two derivations from the same state, but they run at different times, and nothing in between binds them together. If any of the following happens between "approve" and "execute" the bypass lands:
- The state gets edited by another workflow branch, a memory update, or an incoming message.
- The approval payload is a summary rather than the canonical action, and the executor uses a different field.
- The tool arguments are recomputed from context that the reviewer did not see.
The paper's proposed mitigation is short: render the canonical action the executor will use, sign the exact bytes, and compare them at execution time. Or prevent unauthorized pending-state mutation between approval and use. Nothing exotic — but neither is currently the default in the three runtimes named above.
This is what everyone is converging on
Loopjacking landed one day before APort Vault, a benchmark we wrote up yesterday that ran 4,371 attacks against 14 models acting as payment agents. Model-only defense failed 74.6% of the time. A deterministic pre-action check bound to the exact action drove that to 0% across 879 attempts. Same reading as Loopjacking, arrived at from a different angle: the boundary has to be enforced from outside the model, on the actual bytes about to fire, verified against something that cannot mutate after the fact.
Spain's AEPD got there before either paper. Its February 2026 agentic-AI framework — cited as the basis for a September breach notification — includes the "Rule of 2": an agent must never simultaneously process untrusted input, access sensitive data, and take autonomous action without oversight. September gave the regulator its first matching Article 33/34 filing.
Two papers, one regulator, three independent framings. All pointing at the same failure: the last boundary — human approval — is only a boundary if it is bound to the exact action.
What defenders should do this week
Three concrete moves, none of which require the paper's specific engineering:
- Check for the named versions. If you're running Agno AgentOS ≤ 3.0.9, LangGraph Agent Server ≤ 0.14.0 in a conditional in-memory composition, or OpenClaw ≤ 2026.2.23, you are running a runtime with a reproduced attack against your approval gate. Upgrade where a fix ships; segment where it does not.
- Audit what your approval UI actually shows. If the reviewer reads a natural-language summary and the executor consumes tool arguments recomputed at run time, you are in the representation-mismatch shape whether or not you use one of the named runtimes.
- Assume post-approval state can be mutated. Any pending-state store the workflow reads between approval and execution is inside the attack surface. Sign the canonical action at approval time; verify at execution time; refuse if they differ.
None of that is bespoke to any one product. All three papers point at it. The reason the industry hasn't shipped it yet is that "the approval prompt fires" felt like enough.
Sources
- arXiv Sept 21 — Loopjacking: Hijacking Human-in-the-Loop Approval
- arXiv Sept 22 — APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
- SPR{K}3 blog — A Payment Agent Benchmark Just Ran 4,371 Attacks. The Model-Only Defense Loses Three Out of Four Times.
- BleepingComputer — Spain's data agency gets first report of AI-powered data breach
- Agentic Security Newsletter — Week of September 21, 2026
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.