Two LightLLM Helper Ports Expose Code Execution on Inference Nodes
The model API is the obvious place to put authentication and network controls. Two LightLLM flaws disclosed on 29 September concern different listeners on the same inference nodes. Researchers found that an unauthenticated network peer could send a crafted object to an internal RPyC service and make the node run Python code with the LightLLM process's privileges. The reports cover LightLLM through version 1.2.0. They do not establish that anyone has exploited either flaw outside the researchers' tests.
The first path, CVE-2026-103041, is the embed-cache service used by multimodal deployments. LightLLM can enable that service automatically for vision and audio models. The second, CVE-2026-103040, is a router profiling service started when an operator enables profiling. Both reports describe a service bound to all network interfaces without authentication, with Python pickle deserialisation permitted. Each listens on a port allocated at startup, separate from the normal model API port.
That distinction is the operational point. A gateway in front of the model API does not necessarily guard these helper ports. Neither reported path requires a hostile model file, a prompt injection or a valid API token. The attacker does need network reachability to the particular RPyC listener. A node behind a correctly enforced network boundary may therefore be insulated from this route, while a node reachable by other workloads in a shared cluster may remain exposed.
Check the listeners, then the boundary
Start with an inventory of deployed LightLLM versions and model types. Flag vision and audio deployments for the embed-cache check, and find any node launched with profiling enabled. On each affected node, inspect the actual listening sockets and LightLLM startup logs rather than assuming that the documented model API port is the only ingress. Record which source networks can reach each RPyC port. The researchers reproduced both routes from a second host, so a test from another workload in the same network segment is more useful than a local-only check.
Restrict the helper listeners at the host firewall or network policy so only components that need them can connect. If profiling is not required, turn it off. For multimodal workloads, isolate the inference nodes until you can enforce and verify the port boundary. Keep the intended model API available through its normal controls, and confirm from an untrusted peer that connections to the helper ports fail. A successful health check is not a safety test: in the published reproductions, the node kept serving requests after the test payload ran.
At our 30 September review, both original issue reports remained open, and we could not verify a release that fixes these two paths. Do not treat an unspecified upgrade as remediation. Track the maintainer's response, verify the exact code and build deployed when a repair arrives, and retain network restrictions in the meantime. If an unexpected peer could already reach a helper port, investigate connection history and the permissions and secrets available to the LightLLM service account. The public reports show a route to code execution, not evidence of a breach in any particular deployment.


