The Command-and-Control Server Is Now a Panel of Four Chatbots
Cisco Talos published a Windows credential stealer on Monday that does not have a command-and-control server. Every 5 to 15 minutes it sends the host's computer name, its Windows version, and whether the current user is an admin to four different commercial LLM providers — DeepSeek, Qwen, Mistral, and Google Gemini — asks each one to pick from a short list of actions, counts the votes, and executes what the majority picks. The winning action, and every model's stated reasoning for how it voted, gets shipped to the attacker through a Discord webhook.
Talos calls the family CLOSEDQUORUM and, in the primary writeup, frames it as "to our knowledge, the first publicly documented Windows implant" to delegate command-and-control decisions to a panel of AI models. BleepingComputer, SiliconAngle, Unite.AI, and Techzine all picked it up the same day.
The sample is not theoretical. Talos has six SHA-256-hashed builds covering roughly a week of the developer's build chain; the earliest code is dated June 17, 2026. The distributed sample carries placeholder API keys and a placeholder Discord webhook — nobody's confirmed live infections yet — but the code paths are real and the developer's forum posts about carding go back to 2025.
What "the model votes" actually means
The malware ships four candidate actions to the panel: steal, inject, persist, move. Per Talos, each model has to respond in a specified schema — if the reply doesn't parse, that vote is discarded. Ties are broken by a fixed provider order: DeepSeek first, then Qwen, then Mistral, then Gemini. The steal handler pulls Windows credentials out of process memory, saved passwords out of Chrome, Edge and Firefox, and wallet files out of MetaMask, Exodus, and Ethereum keystores. The inject handler generates shellcode and injects it into another process. The persist handler drops a WindowsUpdate value under the user's Run key and adds a WMI event with a Windows-Update-themed name. The move handler — lateral movement — has no implementation in the analyzed build; that one appears to still be on the developer's roadmap.
None of those steps require an AI model to work. What the panel decides is which one runs next.
Why the C2 server is gone
Talos names the substrate reframe out loud. The quote from the primary: "Traditional C2 architecture requires the attacker to operate server infrastructure: a domain, an IP, a protocol, and a listener. That infrastructure is attributable, blockable, and expensive to rotate. … Instead of a singular, unique C2 server, CLOSEDQUORUM calls up to four commercial LLM provider endpoints used by thousands of legitimate applications daily."
That is the whole reason the design exists. Every defender playbook for taking down a C2 infrastructure starts with the same primitives: sinkhole the domain, block the IP range, drop the URL from egress, share the indicator with your ISAC. All four levers assume the C2 endpoint is unique to the attacker. When the endpoint is api.deepseek.com — or the corresponding Qwen, Mistral, or Gemini URL — the endpoint is the same one every legitimate AI application in the environment is already calling. Blocking it takes down the malware and also takes down every legitimate AI feature at the same time. Sinkholing is not on the table. Attribution collapses because there is nothing attacker-controlled to attribute.
The attacker's own server is a Discord webhook, which is at least block-listable in principle — but it only carries stolen data and the panel's reasoning trace outbound. The decision loop keeps running against the four LLM endpoints whether the Discord side reaches its destination or not.
The panel is the point
A single model being wired into malware is not new. Anthropic's Sept 11 threat report named GTG-20006, a Russian state-actor cluster using Claude in an auto-rebuild loop. Sysdig's Jul 2026 LLMjacking-Evolved writeup described VAPT, an attacker framework that wires a hijacked Ollama server into a multi-stage offensive pipeline. Sophos documented an "AI agent lab" for EDR evasion that uses local inference to rewrite its own commands at runtime. All of those use one model.
CLOSEDQUORUM uses four, in parallel, from four independent providers. That is not incremental. It buys the attacker three things at once.
The composite intent is invisible to any one provider. DeepSeek sees a short JSON asking it to vote between steal, inject, persist, and move — with no context that its vote is one of four, that the caller is malware, or that other providers are being asked the same question. Same for Qwen, Mistral, and Gemini. Every provider's abuse-monitoring runs on what its own API sees, and what its own API sees does not add up to a hostile decision loop. The panel-level intent lives only at the malware, and — on the observer side — at the operator's Discord channel.
Provider outages don't stop it. If any one of the four throttles or refuses, the malware still gets a plurality from the other three. Rate limits, API-key revocations, and abuse-driven access removals each subtract one vote, not the whole loop.
The operator sees the models' reasoning. The Discord channel receives every model's per-vote text, not just the winning action. That is a shipped-with-the-malware feedback loop for the attacker to tune the prompt over subsequent builds. The developer already had six builds in a week; this is the tuning surface.
The observability shift
The important thing about CLOSEDQUORUM is not the specific implementation. It is that the network primitive defenders have relied on since 2003 — a distinct, attacker-controlled endpoint the attacker's malware phones home to — has been substituted with a fabric of shared commercial API endpoints that every legitimate AI application in the same environment is already using.
The runtime observable that catches it is not the network address the malware talks to. It is the pattern of calls: a local process, every 5 to 15 minutes, calling four different LLM providers with a schema-constrained action-selection prompt, and then acting on the modal reply. That pattern does not have a benign explanation. Legitimate AI applications call one provider, not four in parallel; they call in response to user input, not on a fixed 5-15 minute cadence; they consume the reply, they do not execute a strategy chosen by majority vote across independent providers.
Talos shipped CAIRN — a metadata-first open-source hunting toolkit — the same day, on GitHub, with 24 acquisition filters (provider-api-integration, python-ai-scripts, ai-analysis-evasion, local-llm-runtime, agentic-tooling, and 19 more), a three-tier YARA-based ontology (T1 primitive artifacts, T2 behavioral context, T3 operational families), and CLOSEDQUORUM and LAMEHUG as its two seed reference families. That is a defender-side recognition that AI-integrated malware is now its own hunting category — worth its own ontology and its own hunting language — rather than a novelty class inside a bigger malware corpus.
The takeaway
Two years of defender writing about AI-integrated attacks has framed the model as the thing being tricked: the prompt-injection victim, the poisoned classifier, the agent that was social-engineered. CLOSEDQUORUM is one of the cleaner statements of the other half of the picture: the model as the attacker's own decision engine, distributed across four independent providers so no one provider's abuse detection sees the whole task. The endpoint that used to be evil-domain[.]tld is now a POST to api.deepseek.com, and three others just like it.
That reframes what "block the C2" means. It also reframes what a defender has to look at. The runtime observable is not the network address — it is the shape of the calls a local process is making to shared, ordinary AI infrastructure. That has to be observable from outside the model, on the substrate the malware is running against, because the model itself will always answer the prompt.
Sources
- Cisco Talos — The Closed Quorum: Inside the first reported autonomous AI C2 implant
- Cisco Talos — Introducing CAIRN: Frontier tracking for AI-integrated malware
- CAIRN source — github.com/Cisco-Talos/CAIRN
- BleepingComputer — New ClosedQuorum Windows malware uses AI for attack decisions
- SiliconAngle — Cisco Talos finds malware that puts its next move to a four-model vote
- Unite.AI — Cisco Talos Open-Sources CAIRN to Hunt AI-Integrated Malware
- Techzine — Cisco CAIRN detects AI malware without opening the binary
- GBHackers — Cisco Talos Launches CAIRN Tool to Hunt and Track AI-Integrated Malware
- Sysdig — LLMjacking Evolved: Stolen AI Compute as Offensive Infrastructure
- Help Net Security — Sophos uncovers AI-powered malware lab built for EDR evasion
- Anthropic — Threat Report, August 2026
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.