The Safety Filter and the Shell Are Reading Two Different Commands

July 3, 2026 · SPR{K}3 Research

Ten of the eleven most popular open-source AI coding agents will refuse to run rm -rf ~. Ask them to run r''m -rf ~ instead, and most of them wave it through — then Bash strips the empty quotes and deletes your home directory anyway. The filter and the shell were never looking at the same command.

That gap has a name now: GuardFall, published by Adversa AI and covered by The Hacker News on June 30, 2026.

Ten out of eleven

Adversa tested eleven open-source coding and computer-use agents — a group that together carries roughly 548,000 GitHub stars. Ten of them shared the same structural flaw: opencode, Goose, Cline, Roo-Code, Aider, Plandex, Open Interpreter, OpenHands, SWE-agent, and Hermes, where the bug was first reported in the project's own issue tracker. Only one, Continue, was built to defend against it.

The trick isn't new. Quoting tricks that survive a naive text filter but vanish once Bash parses them have been public knowledge for decades. What's new is who's running the shell now: an AI agent, with your full account access, executing commands an attacker planted in a file it read — a build script, a piece of "documentation," a config shipped inside a cloned repository.

Why a blocklist can't see it coming

Most of these agents try to stay safe the same way: check the command text against a list of dangerous patterns before running it. The flaw is that the filter reads the command as a string, while Bash rewrites that string before executing it. Bash strips quotes, expands variables, and resolves aliases — none of which the filter accounts for. A pattern-matcher looking for rm sees nothing wrong with r''m, because as plain text those are different strings. Bash removes the empty quotes and runs rm regardless.

The same idea generalizes: a destructive command hidden in base64 and piped into a shell, or an everyday tool like find or dd turned harmful with the right flag. Adversa doesn't call this a single bug — it's calling it "a dangerous convention and a class of problems," which is also why there's no single CVE to patch. Adding more entries to the blocklist doesn't close the gap; the mismatch between what the filter reads and what the shell runs is structural.

Two things have to line up for the attack to land, and Adversa notes neither is exotic: the model has to produce the disguised command (easy — tuck it into something that looks like routine work, not a bare rm -rf request), and the agent has to be running unattended, with an auto-execute flag on or its sandbox switched off — both increasingly the default in automated coding pipelines. Adversa demonstrated the full chain end-to-end against the production Plandex binary, using Claude Sonnet 4.6 as the driving model, and reports the same shape worked against eight of the other tools. The firm describes this as lab research; no in-the-wild exploitation has been reported.

The one that held

Continue, the exception, defends by parsing the command the way Bash actually will before deciding whether to allow it — breaking it into the pieces the shell would produce, checking what would truly execute, and keeping a hard block on the most destructive primitives regardless of disguise. That approach held against every payload in Continue's default editor mode; its command-line auto-run mode did somewhat worse, though the hardest blocks still caught the worst commands. Adversa calls the design portable, estimating a re-implementation is roughly a two-day job for an experienced engineer — which is a useful data point for anyone maintaining one of the other ten tools.

What this fits into

GuardFall lands in the middle of a run of related findings from the same lab this year, including TrustFall against Claude Code, Cursor, Gemini CLI, and Copilot CLI, and a separate deny-rule bypass in Claude Code. It also sits alongside AutoJack and Agentjacking, where poisoned content an agent reads gets turned into commands the agent runs with its owner's privileges. The common thread across all of them: untrusted text keeps reaching a real shell before anything understands what that shell will actually do with it.

Until the agents themselves catch up, Adversa's practical advice is worth repeating: point $HOME at a throwaway directory when running an agent unattended, turn off auto-execute flags unless a human genuinely can't be in the loop, never let an agent run on pull requests from forks, and treat any config file shipped inside a repository as untrusted code, since a malicious one can trigger the whole chain on the first accepted edit.

This is exactly the kind of risk that motivated us to build Defend around behavior instead of text patterns — because a blocklist that reads commands one way while the shell reads them another way will always have a gap in the middle, no matter how many patterns you add to it.


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.