Threat
Unpatched Critical LMCache Flaw Lets Unauthenticated Attackers Run Code Remotely
First reported thehackernews.com
Page published · Page updated
Earliest dated coverage: 7 Oct 2026 · First observed: 7 Oct 2026 · Latest dated coverage: 7 Oct 2026
Coverage timeline
Single-source advisory — one report is available.
Why it matters
LMCache is widely deployed infrastructure for speeding up LLM inference, so an unpatched unauthenticated RCE exposes AI-serving backends to full compromise by a single network message.
A critical vulnerability (CVE-2026-105192) in LMCache, open-source software that accelerates LLM servers such as vLLM, allows unauthenticated attackers to execute code remotely via pickle deserialization over the ZeroMQ transport in its multiprocess mode. JFrog disclosed the flaw on October 7, and no fixed version is available; the cache server is exploitable over the network when an operator configures it to listen on a routable address rather than the default localhost.
Summary
JFrog disclosed a critical, unpatched vulnerability (CVE-2026-105192, CVSS 9.8) in LMCache, open-source software that accelerates LLM servers such as vLLM. The flaw lets an unauthenticated attacker run code on the LMCache cache server via pickle deserialization on the multiprocess ZeroMQ transport.[1][8]
The vulnerability affects LMCache from version 0.3.9 (October 2025) through 0.5.5 (the latest stable release), as well as the 0.5.6 release candidates and the development branch. No fixed release exists, and LMCache has not published a security advisory. The report describes a disclosed vulnerability with no evidence of in-the-wild exploitation.[1][12]
Exploitation requires the multiprocess server to be bound to a routable address rather than its default localhost binding; LMCache's own example Kubernetes deployment listens on every interface, and on official container images the process runs as root, elevating impact.[1][10]
Attack chain
- Access: The multiprocess cache server is reachable over the network only when an operator starts it with a routable address instead of the default localhost binding; LMCache's example Kubernetes deployment listens on every network interface.[1][10]
- Exploitation: An attacker sends a single crafted network message to the unauthenticated ZeroMQ socket. The server unpacks one message type with pickle while reading the message's arguments, before any check of the message type, so the attacker's code runs as the data is decoded.[1][11]
- Code execution / impact: The injected code runs with the privileges of the LMCache process, which is root on the project's official container images.[1]
Disclosure timeline
| Date | Event |
|---|---|
| October 2025 | LMCache version 0.3.9, the first affected version, is released.[1] |
| September 22, 2026 | vLLM 0.30.0 is released, fixing the related denial-of-service flaw CVE-2026-105756 affecting deployments using the LMCache multiprocess connector.[1][16] |
| October 6, 2026 | A GitHub user opens six additional LMCache security reports alleging unauthenticated access to cached tenant data and network services that execute commands without a login.[1][13][14] |
| October 7, 2026 | JFrog discloses CVE-2026-105192 with a CVSS score of 9.8; no fixed version is available.[1][8] |
How it works
LMCache's multiprocess mode runs the cache as a standalone server that LLM worker processes reach over the ZeroMQ messaging library. The ZeroMQ socket opened for workers to register and share cached data has no authentication.[1]
One type of message is unpacked with Python's pickle, a format that can carry and execute code as data is decoded. The server unpacks the message while still reading its arguments, before any check of the message's type, so a crafted message runs the sender's code with the privileges of the LMCache process — root on official container images.[1][11]
Network reachability depends on configuration: by default the server listens only on localhost, but it becomes remotely exploitable when an operator binds it to a routable address, as LMCache's example Kubernetes deployment does by listening on every interface. A firewall limiting access lowers but does not remove the risk, since any host that can open a connection can run code.[1][10]
Affected versions and patch status
| Product | Affected | Patch status |
|---|---|---|
| LMCache | Versions 0.3.9 (October 2025) through 0.5.5 (latest stable), the 0.5.6 release candidates, and the development branch | No fixed version; LMCache has not published a security advisory[1][12] |
| vLLM (related flaw CVE-2026-105756) | Versions before 0.30.0 using the LMCache multiprocess connector (denial of service via malformed cache_salt, rated 6.5, no code execution) | Fixed in vLLM 0.30.0, released September 22, 2026[1][16] |
Key takeaways
- CVE-2026-105192 is an unauthenticated RCE in LMCache with no available patch and no published vendor advisory, so operators must rely on configuration-based mitigations.[1][12]
- Impact is severe because official container images run the LMCache process as root, and exposure hinges on a single binding setting that LMCache's own example deployment sets to all interfaces.[1][10]
- The root cause — passing data from an unauthenticated network socket to pickle — repeats the ShadowMQ class of flaws found across AI inference frameworks in November 2025, though a shared code source for LMCache has not been established.[1]
- JFrog's advisory provides no way for operators to determine whether a server has already been attacked, complicating incident response.[1]
Defensive actions
- Do not assign the LMCache multiprocess server a routable address; keep its port bound to the local machine or a trusted cluster network.: JFrog advises this until a patched release ships; the server is remotely exploitable only when bound to a routable address rather than the default localhost.[1]
- Use firewall rules to limit which hosts can reach the multiprocess server port, while understanding this is only a partial mitigation.: A firewall lowers risk but does not remove it, because any host that can still open a connection can run code.[1]
- Review LMCache deployment manifests, particularly the project's example Kubernetes DaemonSet, for servers bound to all network interfaces.: LMCache's own example Kubernetes deployment starts the server listening on every network interface, exposing it to the network.[1][10]