Plugin4Shell: Four Coding Agents, One Wrong Assumption, and Two Vendors That Won't Fix It
Pin the plugin. Name the commit. Review the code at that commit. Ship. The agent will refuse to install anything but that reviewed snapshot — that is what the pinning is for. Except the agent doesn't check whether the code it got actually matches the hash it named. It checks the fetch and moves on. So the person who controls the plugin's repository swaps the code out from under the pin. Background auto-update ships that swap to every installed copy. No user click. No approval prompt. Attacker code, running as the developer, on machines that never opened the agent this week.
That is Plugin4Shell, disclosed Sept 17 by AIR Security researchers Or Nevo, Dor Granat, and Niv Hoffman. Four products are affected — Anthropic Claude Code, OpenAI Codex, GitHub Copilot, and Google's Gemini CLI. AIR reported to all four vendors in June after finding it in May with working exploits. Per Help Net Security, Anthropic shipped a fix in a mid-September Claude Code release and OpenAI shipped one in Codex 0.146.0. Copilot has no fix. Google will not patch the Gemini CLI, which it is retiring.
What the shared assumption is
AIR's writeup is the tightest per-primitive description. Naming a commit SHA and verifying it look like the same operation. They are not. The agents fetch the pin. They check the fetch. They do not check that the bytes they received hash to the value they asked for. If the marketplace returns different bytes under the same pin, the agents install them and treat them as the reviewed code.
Per InfoWorld, the plugin code inherits the developer's whole environment: local source, cloud credentials, SSH keys, access to internal repos and production systems. That is the CVSS 10.0 shape — one primitive, one action, reach equals whatever the developer has. Per Cybersecurity News, both patched agents ship background plugin auto-update as default. A marketplace-side SHA bump doesn't need a running agent session. The plugin the developer installed weeks ago updates itself in the background, and that update is the attacker's code.
Why four vendors shipped the same bug
The plumbing under a plugin marketplace looks like a solved problem — Git has content-addressed storage, package managers have integrity manifests, container registries have signed digests. Each one names the artifact by hash and checks the artifact you got matches. When the AI-coding-agent products layered marketplaces on top, they picked up the naming and left the checking out. Four independently built products landed on the same partial answer.
This is the pattern this quarter. September's BragJack disclosure showed five agentic browsers shipping the same extension-to-agent trust-boundary failure. Docker Sandboxes CVE-2026-77179 showed the microVM built to isolate coding agents shipping a symlink-race guest-escape. The MervinPraison PraisonAI cluster shipped five CVSS 9.8 unauth-RCE issues in one repo in one day. Plugin4Shell is the same shape on the plugin marketplace: an integrity primitive the ecosystem thinks is in place, isn't.
What two-fixed / two-not says about the market
Anthropic and OpenAI shipped patches. GitHub Copilot has not, per Forkast. Google's response, per AiCybr: the Gemini CLI is being retired and will not be patched. The final shipping version carries the flaw. Developers who installed a Gemini CLI plugin during 2026 and stopped thinking about it are left holding the risk.
A patch is the vendor's answer to a finding. Product-sunset without a patch is the vendor answering that they aren't going to answer, and users who trusted the pin don't know until they read the news cycle. Runtime observability — the layer that watches what the developer's process does with a plugin, and reacts when a background update turns it into something the developer never approved — keeps working when the vendor's answer is "we're not going to answer." Same layer keeps working after the next four vendors ship the same bug.
Where the check has to live
Every AI-coding-agent brand now ships a plugin marketplace. Each has to answer the same question the agents got wrong: does the code the agent ran actually match the code the developer reviewed? A per-vendor answer in a per-vendor plugin loader works until the next vendor. An answer written outside the agent — watching what the developer's process does when a plugin loads, noticing when a background auto-update executes code the developer never approved, treating outbound source or credential push from a plugin refresh as an event that has to look like something the developer did — is the layer that survives another four vendors shipping the same failure. Plugin4Shell, BragJack, and the PraisonAI MCP-connect CVEs all land on that layer.
Sources
- The Hacker News — Plugin4Shell Lets Repository Owners Swap Pinned Plugin Code Across Four AI Coding Agents
- Help Net Security — Zero-click RCE vulnerability hit four major AI coding agents, two remain unpatched
- AIR Security — Plugin4Shell: Zero Click RCE Vulnerability found in top 4 most popular coding agents
- InfoWorld — A zero-click RCE flaw in AI coding agents could have exposed enterprise systems
- Cybersecurity News — Plugin4Shell Zero-Click RCE Hits Claude Code, Codex, Copilot and Gemini CLI
- FourWeekMBA — Claude Code and Codex Patched After Plugin4Shell Exposes a Shared Assumption Across Four AI Agents
- Forkast — Plugin4Shell Bypasses SHA Pinning Across All Four Major AI Coding Agents
- AiCybr — Plugin4Shell: Claude Code and Codex Patched, Copilot and Gemini CLI Remain Exposed
- The Hacker News — Critical Docker Sandboxes Flaw Lets Malicious Guest Code Read and Modify macOS Host Files
- The Hacker News — One Extension Could Hijack AI Assistants Across Chrome, Comet, Edge, Opera Neon and Claude
- Strix.ai — CVE-2026-57123 (PraisonAI cluster)
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.