A Payment Agent Benchmark Just Ran 4,371 Attacks. The Model-Only Defense Loses Three Out of Four Times.

September 22, 2026 · SPR{K}3 Research

Point a tool-using AI agent at a task that ends in "send this payment" and let a room full of humans try to trick it. Then run the same attacks again, but this time make the last step check the intended action against a signed policy, outside the model, before the money moves. In the first setup, the attackers win about three times out of four. In the second, they win zero times out of 879. That is the APort Vault benchmark, released Sept 22 alongside a 225,964-row public dataset on Hugging Face.

The setup

Between March and August 2026, a public capture-the-flag event collected 4,371 human-authored attacks against a live payment agent. APort's team replayed those attacks against 14 different frontier models from eight labs, in five policy configurations, on two tracks: one where the model alone decides whether the payment is legitimate, one where a deterministic pre-action check — the Open Agent Passport — compares the actual pending action against a signed authorization scope and blocks anything out of bounds. Same attacks, same models, same task. Two answers.

What the numbers say

Under model-only defense, 74.6% of social-engineering attacks succeeded in getting the agent to pay the attacker. Under the OAP pre-action check, unauthorized transfers hit 0% across 879 attempts. That gap does not depend on the model. It reproduces across all 14 tested models from all eight labs. Better instruction-tuning, longer safety training, more careful system prompts — none of that closed the gap. The gap is architectural.

Why the model can't be the boundary

An LLM handling a payment sees the same context window whether the request is legitimate or a social-engineering attack. There is no runtime separation between the trusted instruction ("only pay invoices this operator has approved") and the untrusted content ("urgent, CFO says wire this now, here is the reason"). Prompt injection researchers have been saying this for two years — OWASP's 2026 LLM Top 10 still lists prompt injection as risk #1, and Help Net Security reported this June that prompt injection remains the failure mode driving most agentic-AI security incidents in production. APort Vault is what that abstraction looks like at 4,371 attacks and a real payment endpoint.

The paper's authors are direct about it: controls that sit outside the model, bound to the exact action being executed, hold up. A loosely coupled approval prompt handed back to the model does not.

The Loopjacking result cuts the same way

APort Vault landed one day after a separate arXiv paper on the same shape of problem. Loopjacking, published Sept 21, shows that even when a human is looped into the approval — the last, hardest boundary anyone actually deploys — an attacker can arrange for the human to approve operation A while the workflow later executes materially different operation B. The Loopjacking authors reproduced this in seven Agno AgentOS releases ending at 3.0.9, twelve LangGraph Agent Server versions ending at 0.14.0, and OpenClaw 2026.2.23. The approval prompt fires. It just shows the wrong operation, or the workflow state gets mutated between approval and execution.

Two independent September papers, converging on one reading: the boundary has to be enforced from outside the model, on the actual action about to fire, verified against a policy the workflow cannot mutate after the fact.

Not new — now measurable

Spain's data-protection regulator arrived at the same rule from a completely different direction. AEPD's February 2026 agentic-AI framework coined a "Rule of 2": an agent must never simultaneously process untrusted input, access sensitive data, and take autonomous action without oversight. Spain got its first matching Article 33/34 breach notification in September — an unnamed organization's agent scanned public files, exploited a vulnerability, modified personal data, and reached invoice records without a human directing each step. The regulator, the payment benchmark, and the human-in-the-loop paper are pointing at the same failure surface.

What this means for anyone shipping an agent that touches money, records, or infrastructure

The industry defaults, right now, are:

APort Vault says the first default fails 74.6% of the time against a public attack corpus. Loopjacking says the second is bypassable when the workflow state is mutable. The remediation both papers point at is the same: a deterministic check, outside the model, bound to the exact action, verified against a policy signed before the attack started.

That's the class of defense to build against, whether you ship it yourself or buy it. The one thing you cannot do is nothing — the attackers now have a 4,371-row public corpus to train against.

Sources


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.