Prompt Injection Is 200 Patterns and Counting
On July 7, CrowdStrike updated its public prompt-injection taxonomy with 18 new patterns. The total crossed 200. Two years ago, when the first OWASP LLM Top 10 shipped, the whole category was a single-paragraph entry. That number, and the shape of the curve it's on, tell you something specific about what defense looks like from here.
What the taxonomy actually covers
CrowdStrike's taxonomy is not a list of specific attacks. It's a list of patterns — the reusable techniques an attacker composes to get a model or an agent to do the wrong thing. Two of the new July additions are worth naming, because they aren't hypothetical:
Trigger-Activated Rule Addition is the class where an attacker plants instructions that sit dormant in an agent's context until a keyword or condition activates them. Last week's Ghostcommit disclosure — the PNG-hidden prompt injection that fooled CodeRabbit and Cursor Bugbot — is the pattern's cleanest live example. The AGENTS.md file in the merged PR is the dormant rule. The trigger is "a developer, in an unrelated session days later, asks the agent for a routine feature." The exfiltration procedure lives inside a PNG the review layer structurally cannot see.
Cognitive Token Suppression works one layer down. Instead of trying to bypass the safety training that produced a refusal, it steers the model's linguistic choices away from the specific words the refusal comes wrapped in. The model doesn't hit the phrases it was trained to output on this class of request. It also doesn't hit the tripwire that flags them.
These are two different attack primitives. They compose with other primitives already in the taxonomy — role spoofing, style injection, multimodal authority markers, indirect retrieval — into thousands of specific attacks. The taxonomy is 200 patterns. The combinatorial space is much larger.
The shape of the curve is the point
Any single technique on the list is patchable. Vendors have patched dozens of them. The Cursor sandbox got a fix. Anthropic added symlink resolution to Claude Code. Google patched Antigravity. Writer isolated the sandbox origin. Each patch closes a specific address on a specific product.
The taxonomy has been growing all year. The patches are keeping pace with individual attacks and losing ground on the category. That's a specific kind of curve, and it doesn't turn around by trying harder. Every new modality — vision, audio, code, tool-use — opens new pattern space. Every new agent framework (MCP, A2A, coding CLIs) adds new content-to-instruction boundaries the model has to guess about. The number goes up because the surface goes up.
The MIT Role Confusion paper (Ye, Cui, Hadfield-Menell) gives the deeper reason: the model can't reliably tell the difference between "content" and "instruction" because the boundary lives in writing style, not in the API-level role tag. Anything that produces the right style — a system-prompt header rendered inside an image, a SYSTEM:-styled block in retrieved web content, an AGENTS.md in a merged PR — reads as authoritative. There is no fixed set of tokens the model can be trained to always refuse. That is what "200 patterns and counting" means in practice.
What that implies about defense
Two things follow, and both of them are structural.
Per-model patching cannot close the category. Not because vendors are slow, but because the boundary that fails is not one the model has representation for. You can retrain against every known pattern. The next paper adds a new one. This is why the vendors most invested in RLHF-style safety training are the ones publishing the largest taxonomies of things RLHF doesn't fix. CrowdStrike isn't the outlier; it's the current data point.
The durable defense has to live outside the model. If you can't teach the model to distinguish instructions from data, you have to watch what the model tries to do, compare it to what the user actually authorized, and refuse the actions that don't line up. That's the whole content of Ghostcommit's counterexample: the Claude Sonnet model that leaked .env files under Cursor and Antigravity refused the same request when run under Claude Code. The weights were identical. What changed was the scaffold around them — what content in context got treated as an instruction and what got treated as data. The boundary that held was the runtime boundary, not the model boundary.
What we take from this
If you're defending an AI agent this quarter, the shape of the taxonomy curve is your planning input. Two hundred patterns means the roadmap of "patch each one as it lands" ends where the attackers' publication rate begins. The math is not on your side.
Runtime behavioral monitoring is where the math flips. There are many ways to write a prompt injection. There are far fewer ways for a compromised agent to cause damage — write to ~/.ssh, read .env, exfiltrate over the network, install a persistent daemon. Watching those exits catches the entire pattern space at the point where all of it converges. That is where the durable defense lives.
Sources
- CrowdStrike — CrowdStrike Uncovers New Prompt Injection Techniques
- CrowdStrike — Taxonomy of Prompt Injection Methods poster
- ASSET Research Group — Ghostcommit disclosure
- BleepingComputer — Ghostcommit hides prompt injection in images to fool AI agents, steal secrets
- Ye, Cui, Hadfield-Menell — Prompt Injection as Role Confusion (arXiv:2603.12277)
- OWASP — Top 10 for Large Language Model Applications
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.