One Network Message, Root on Your LLM Node
JFrog's security research team disclosed CVE-2026-105192 yesterday. CVSS 9.8. Unauthenticated remote code execution in LMCache, the open-source cache layer that sits in front of vLLM in a lot of production LLM-serving stacks. No patch at publication. The official container runs the vulnerable process as root.
The bug is a two-sentence story. LMCache's multiprocess server exposes a ZeroMQ socket. The socket has no authentication, and the server calls pickle.loads on the inbound bytes before it checks what kind of message they are. One crafted network message gets you code execution as the LMCache user. On the published Kubernetes example, that user is root.
Where it ships
LMCache is not a research toy. It is the KV-cache tier used behind vLLM to keep repeated prefixes off the GPU, and vLLM is the most-deployed open-source LLM serving framework we have. The affected versions, per JFrog's advisory, are 0.3.9 (released October 2025) through 0.5.5 (the latest stable), plus the 0.5.6 release candidates and the development branch. A straight year of releases.
By default LMCache binds the socket to localhost. The teams that got exposed are the ones who followed LMCache's own example Kubernetes DaemonSet, which listens on all interfaces so that worker pods elsewhere in the cluster can reach it. That is the standard pattern for a cache tier in a multi-worker LLM deployment. If the quickstart was followed as written, every pod on the cluster's network can send one ZeroMQ message to that socket and get code execution.
The container images that LMCache publishes run the process as root. The exploit lands with root inside the LLM worker's container.
Why pickle keeps winning
The specific primitive here — call pickle.loads on bytes that came off the network, decide what type of message it was afterward — is the same thing HiddenLayer and Microsoft documented in November 2025 under the "ShadowMQ" label across a cohort of AI-inference frameworks. JFrog's own advisory draws the line explicitly and says whether LMCache shares code with the ShadowMQ cohort "has not been established."
Eleven months. One well-documented pattern. Still shipping, with root containers, in the cache tier of the most popular open-source LLM serving framework. Pickle is a Python-native serializer that was never meant to receive untrusted input; its documentation has said so for decades. The way it keeps slipping back into production code is not that developers forget. It is that the pattern "deserialize first, dispatch on type afterward" is the default shape of an RPC server written in a hurry, and the inference-serving substrate is being written in a hurry.
What the defender can and cannot do
JFrog's recommendations are the right ones for right now: do not bind the LMCache socket to a routable address, keep it on localhost, limit the cluster network segment that can reach it, read the LMCache config before you run the quickstart. There is no fixed version to upgrade to.
Where it stops being enough is where the detection stops working. JFrog's own advisory says there is no way for an operator to tell whether a server has already been exploited. The exploit is a single pickle payload. It leaves no distinguishing log line in LMCache itself. The attacker is in as root with the same identity LMCache would normally have. By the time the fixed version lands, the question of "was this server exploited before we patched it" has no answer from the artifact or the logs.
The surviving signal is behavioral. A cache server does a narrow set of things: receive key-value writes, serve key-value reads, talk to the LLM workers it is attached to. A cache server that spawns a shell, reads credentials, calls out to the network, or writes to the filesystem outside its cache directory is doing something it has never been built to do, regardless of what version of LMCache it is running. That is the layer the defender still has, because the attacker's goal — persistence, lateral movement, exfiltration — forces behavior that a cache process has no business performing.
This is what we built Defend around. The patch tells you what to fix; the behavior tells you whether you were already too late.
The pattern this is part of
LMCache is the second LLM-infrastructure-cache RCE this quarter after the DeepSeek Harness CVE-2026-82533 sandbox escape, and lands into a running cohort that includes the IBM mcp-context-forge CVSS 10 cluster, the Rufroot MCP-Bridge CVE-2026-59726 233-tool compromise, and the Splunk MCP CVE-2026-76404 deserialization RCE. The common shape: the serving substrate itself — not the model, not the agent, not the user's prompt — is the attack surface.
The industry's answer to date has been to add more guardrails inside the agent and more approval UIs around the tool call. CVE-2026-105192 does not care about any of that. It lands behind the LLM worker, before the model even sees a prompt, inside a process that nobody thought about as part of the attack surface because its job was to make inference go faster.
If your LLM-serving stack uses LMCache, two things are worth doing today. Confirm your socket is bound to localhost or a tight CIDR. And start logging behavior at the process boundary of your inference workers and their cache tier, because the single network message that gets exploited leaves no other trace.
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.