Security & privacy

Critical unpatched flaw in LMCache allows remote code execution

A critical vulnerability in LMCache lets attackers run code remotely on vLLM servers. No patch exists yet, requiring immediate network isolation.

Illustration of a vulnerable server with exposed circuits
Illustration generated for this article

Security researchers have identified a critical vulnerability in LMCache, an open-source acceleration layer for large language model servers like vLLM. Disclosed on October 7 by JFrog, the flaw allows unauthenticated attackers to execute arbitrary code on affected systems. As of this writing, no fixed software version is available, leaving operators reliant on network configuration changes for protection.

What happened

The vulnerability, tracked as CVE-2026-105192, carries a severity score of 9.8 out of 10. It affects LMCache versions from 0.3.9, released in October 2025, through the latest stable release 0.5.5. The issue also persists in the 0.5.6 release candidates and the current development branch. JFrog’s security research team, led by Yuval Moravchick, discovered that a single crafted network message can trigger remote code execution on the cache server.

The risk level depends heavily on deployment configuration. By default, the LMCache multiprocess server listens only on the local machine, which prevents external access. However, many multi-node deployments require the server to listen on a routable address to share cached data across machines. LMCache’s own example Kubernetes deployment configures the server to listen on every network interface, effectively exposing it to the cluster network. In these configurations, any host that can reach the port can exploit the flaw.

Compounding the severity, the official container images for LMCache run the process as the root user. This means successful exploitation grants the attacker full administrative privileges on the host. JFrog notes that while firewalls can limit exposure, they do not eliminate the risk entirely if any trusted host within the allowed range is compromised. There is currently no method provided to determine if a server has already been attacked.

How it works

The root cause lies in how LMCache handles inter-process communication using the ZeroMQ messaging library. The multiprocess server opens a socket for worker processes to register and share cached data, but this socket lacks any authentication mechanism. When a message arrives, the server uses Python’s pickle module to deserialize the data. Pickle is known to be unsafe for untrusted data because it can execute arbitrary code during the decoding process.

Crucially, the server unpacks the pickle data while still reading the message arguments, before it checks the message type. This order of operations allows an attacker to send a malicious payload that executes immediately upon receipt. The code runs with the same privileges as the LMCache process, which, as noted, is often root in containerized environments. This pattern mirrors a group of flaws called ShadowMQ, identified in other AI inference frameworks in November 2025, though a direct code link has not been established.

Key details

  • CVE Identifier: CVE-2026-105192, rated 9.8/10 in severity.
  • Affected Versions: LMCache 0.3.9 through 0.5.5, plus 0.5.6 release candidates and dev branches.
  • Attack Vector: Unauthenticated network message via ZeroMQ socket in multiprocess mode.
  • Root Cause: Unsafe deserialization using Python pickle before message type validation.
  • Privilege Level: Code executes as the LMProcess user, which is root in official containers.
  • Patch Status: No fixed version is currently available.

Why it matters

For engineering teams building AI infrastructure, this vulnerability highlights the risks of adopting new acceleration tools without rigorous security auditing. LMCache is designed to speed up LLM inference, a critical performance metric for production applications. However, the default examples provided by the project encourage insecure network bindings. Operators who follow these examples to set up multi-node clusters inadvertently expose their systems to remote code execution.

The absence of a patch forces teams to choose between performance and security. Disabling the routable address breaks multi-node caching, potentially degrading service quality. Keeping it open leaves the door wide open for attackers. This situation underscores the importance of defense-in-depth strategies, such as strict network segmentation and least-privilege principles, rather than relying solely on application-level security.

Additionally, the discovery of six other unconfirmed security reports on GitHub suggests broader hygiene issues within the project. While these lack CVEs or maintainer confirmation, they point to potential unauthenticated access to tenant data and other services. Teams using LMCache must now monitor not just this specific CVE but also the project’s overall security posture and response to emerging threats.

What you can do

  • Restrict Network Access: Configure the LMCache multiprocess server to listen only on localhost or a trusted, isolated cluster network.
  • Avoid Public Binding: Do not bind the server to routable addresses accessible from untrusted networks or the public internet.
  • Implement Firewalls: Use firewall rules to strictly limit which IP addresses can connect to the LMCache port, though note this does not fully mitigate the risk.
  • Check Container Privileges: If possible, modify container configurations to run the LMCache process as a non-root user to limit the impact of exploitation.
  • Monitor for Updates: Closely watch the LMCache repository and JFrog advisories for the release of a patched version.
  • Audit Deployments: Review existing Kubernetes or Docker deployments to ensure they are not using the default example configurations that expose all interfaces.

Tools from the Bytechap store

Keep reading

All stories