RC RANDOM CHAOS

LMCache Exposes Remote Code Execution

CVE-2026-105192 lets unauthenticated attackers run code on exposed LMCache multiprocess servers, with no patched release available.

· 4 min read
LMCache Exposes Remote Code Execution

A single network message to an LMCache multiprocess server can run commands as the user running the cache process, and there is no fixed release available.

JFrog disclosed the flaw on October 7 and rated it 9.8 out of 10 when the server is bound to a routable address. The vulnerability is tracked as CVE-2026-105192. It affects LMCache from version 0.3.9, released in October 2025, through 0.5.5, the latest stable release. It is also present in the 0.5.6 release candidates and the development branch.

LMCache is open-source software used to speed up large language model servers such as vLLM. The vulnerable path is its multiprocess mode. In that mode, the cache runs as a standalone server, and LLM workers communicate with it over ZeroMQ. That server opens a socket used by worker processes to register and share cached data.

The bad part is simple enough to be familiar: the socket has no authentication, and one message type is decoded with Python pickle. Pickle is not a neutral serialization format for untrusted data. It can carry executable behavior, and decoding it can run code. In this case, the server unpacks the message while still reading its arguments, before it checks the message type. A crafted message can reach the pickle decode path and execute code supplied by the sender.

The resulting code runs with the privileges of the LMCache process. JFrog says the process runs as root in the project’s official container images. That turns a reachable cache service into a direct remote code execution target with container-level consequences, and possibly more depending on how the deployment is wired into storage, secrets, scheduling, and model-serving infrastructure.

Exposure depends on one configuration choice. By default, the multiprocess server listens only on localhost, so another host cannot reach it. The service becomes reachable when an operator starts it with a routable address. That is the kind of setting used when multi-node deployments share a cache across machines. LMCache’s own example Kubernetes deployment starts the server this way, listening on every network interface. A copy of LMCache running inside a single vLLM process does not open the port at all.

For operators, the immediate check is whether LMCache is being run in multiprocess mode as a standalone server, and whether that server listens beyond localhost. In Kubernetes, the interesting places are the container command, environment-driven startup flags, Service objects, pod networking, and any NetworkPolicy that decides who can connect. A pod listening on every interface is not automatically internet-exposed, but it is exposed to whatever the cluster network, service mesh, node routing, and firewall rules allow.

JFrog’s mitigation is also direct: do not assign the multiprocess server a routable address until a patched version exists. Keep the port on the local machine or on a trusted cluster network. A firewall limiting access reduces the blast radius, but it does not fix the vulnerability. Any host that can still open a connection to the port can run code.

That distinction matters in inference clusters. The trusted network is often larger than people remember. Batch workers, evaluation jobs, notebooks, model gateways, observability sidecars, and internal test workloads can all end up with paths into services that were never meant to be security boundaries. If LMCache is reachable from other workloads, those workloads are effectively inside the exploit perimeter.

There is also no public guidance from LMCache maintainers at the source described here. LMCache has not published a security advisory for CVE-2026-105192, and JFrog’s advisory does not give operators a way to determine whether a server has already been attacked. That leaves defenders with ordinary incident response work: identify exposed deployments, restrict access, rotate anything the LMCache process could read, and inspect surrounding systems for behavior consistent with code execution from that container or user account.

The adjacent reports are less settled. On October 6, a GitHub user opened six additional LMCache security reports alleging unauthenticated access to cached data across tenants and several unauthenticated command-executing network services. Those reports came from one account, are based on proof-of-concept claims, and have no CVE, maintainer confirmation, or fix. One report points to a default that has changed: an admin HTTP server that listened on every network interface in 0.5.5 listens only on localhost in the 0.5.6 release candidates.

There is one related fixed issue in vLLM. Before version 0.30.0, released September 22, a request with a malformed cache_salt value could crash the engine on deployments using the LMCache multiprocess connector. That denial-of-service flaw is tracked as CVE-2026-105756, rated 6.5, and does not allow code execution.

The engineering failure in CVE-2026-105192 is handing unauthenticated network input to pickle. Researchers found the same class of mistake across other AI inference frameworks in November 2025 in flaws called ShadowMQ, although whether LMCache shares code with those projects has not been established.

Until a fixed LMCache release exists, the useful control is exposure reduction. Find multiprocess LMCache servers, remove routable binds, limit the port to the smallest possible set of trusted peers, and avoid running the process with privileges that make a cache compromise more useful than it already is.

Share

Keep Reading

Latest on the Wire

Full wire →

New signal daily · RSS

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.