The Open Source Cache Layer Under Your LLM Servers Has an Unpatched Critical Flaw With No Fix Date
Cybersecurity

The Open Source Cache Layer Under Your LLM Servers Has an Unpatched Critical Flaw With No Fix Date

JFrog disclosed a 9.8-severity flaw in LMCache, the caching layer many self-hosted LLM deployments use to speed up vLLM, that lets an unauthenticated attacker run code on the server. There is no patch yet.

PublishedOctober 8, 2026
Read time5 min read
Share

What JFrog found

JFrog's security research team, credited to researcher Yuval Moravchick, disclosed CVE-2026-105192 on October 7, a critical flaw in LMCache rated 9.8 out of 10. LMCache is open-source software that accelerates large language model serving frameworks like vLLM by caching key-value data, and in its multiprocess mode it runs that cache as a standalone server that LLM worker processes reach over the ZeroMQ messaging library. The flaw is straightforward once stated: the server unpickles one type of incoming message before it verifies what kind of message it actually received, and Python's pickle deserialization can be made to execute arbitrary code.

Because the ZeroMQ socket carries no authentication, nothing stops an attacker who can reach the port from sending a crafted message directly. The result is unauthenticated remote code execution running as the LMCache process user. On official container images, JFrog notes that user is root, which turns a caching-layer bug into a path to full host compromise. The flaw spans every release from 0.3.9, published in October 2025, through the current stable 0.5.5, and the 0.5.6 release candidates and development branch carry it too.

Why this is exposed by default, not by misconfiguration

The usual caveat with a flaw like this is that exploitation requires an unusual deployment choice, binding a service to a public interface instead of localhost. That caveat does not fully apply here. LMCache's own example Kubernetes DaemonSet configuration, at version 0.5.5, listens on all interfaces, meaning any deployment built directly from the project's published example inherits the exposure without an operator making an explicit mistake. A firewall that restricts which hosts can reach the service lowers the risk, but JFrog is explicit that it does not eliminate it: any host that can establish a connection can still trigger the flaw.

There is a meaningful exception worth noting for scoping your own exposure: a copy of LMCache embedded inside a single vLLM process, rather than run in multiprocess mode as a standalone server, does not open the vulnerable port at all. The risk is concentrated specifically in multiprocess deployments, which are also the configuration most organizations choose when they want to scale LLM inference across multiple workers, which is to say, the configuration production deployments are most likely to use.

The patch gap and what to do instead

LMCache has not published a security advisory and no fixed version exists as of this writing. JFrog's practical guidance is to keep the multiprocess server off routable addresses entirely and restrict it to localhost or a trusted cluster-internal network, which is a network segmentation fix rather than a patch. For any organization running LMCache in production, that means auditing current deployment manifests now, specifically checking whether the service is bound to 0.0.0.0 or a cluster-routable address, as the example DaemonSet configuration does, and remediating that binding immediately rather than waiting for an upstream fix.

This is also a case where a firewall rule is not a substitute for fixing the binding. Because JFrog confirms any host able to connect can still exploit the flaw regardless of access-control-list restrictions on who is allowed to connect, network-layer allowlisting reduces the attack surface but does not close the vulnerability. Teams should also note the related CVE-2026-105756, a lower-severity flaw in vLLM itself affecting versions before 0.30.0 when paired with the LMCache multiprocess connector, which causes an engine crash rather than code execution but confirms the two projects' integration point deserves a broader review.

What this means for self-hosted AI infrastructure decisions

Organizations that chose to self-host LLM inference, for cost, data residency, or customization reasons, took on an infrastructure stack that did not exist two years ago and has not yet accumulated the hardening that mature infrastructure components get over time. LMCache is exactly that kind of component: useful, increasingly common in production deployments, and young enough that a researcher found an unauthenticated RCE in a mode the project's own example configuration ships insecurely. The article links this pattern to ShadowMQ research from November 2025, which found similar pickle-on-unauthenticated-socket flaws across other AI inference frameworks, suggesting this is a class of bug specific to how the AI infrastructure ecosystem adopted ZeroMQ-style messaging without carrying over the authentication assumptions that messaging pattern needs.

That history argues for treating self-hosted AI inference infrastructure with the same security maturity assumptions you would apply to a newly adopted database or message broker, not with the assumption that because it ships with an AI framework it has been hardened the way that framework's model-serving code has been. A deployment choice made for cost or control reasons should come with an explicit security review of every ancillary service in that stack, cache layers very much included, rather than an assumption that the stack's youth means nobody has looked yet.

The roadmap takeaway

If your organization runs vLLM with LMCache in multiprocess mode, check your binding configuration today. If it is reachable from anything other than localhost or a tightly scoped cluster-internal network, that is an unauthenticated root-level remote code execution risk sitting in production right now, with no patch to apply and no timeline for one. The fix is a network configuration change you can make this week, not a vendor release you have to wait for.

Beyond this specific flaw, add LMCache and comparable AI infrastructure components to whatever asset inventory drives your vulnerability management program, if they are not already there. The pace at which new AI serving infrastructure is being adopted is outrunning the pace at which it is being security-reviewed, and the next ShadowMQ-pattern flaw is more likely to show up in a project your team adopted in the last twelve months than in infrastructure that has been through several rounds of security scrutiny already.

Tagged#news#security#cybersecurity#breach#cisa#ransomware#zero-day#supply-chain#ai-security#ai-infrastructure#llm-security#zero-day-rce#lmcache#vllm#jfrog#remote-code-execution#deserialization-attack