What’s trending in AI on 8 October 2026: For two years the AI security conversation has been about what models say and what agents do. This week the attackers went for the plumbing underneath. In 48 hours, researchers disclosed an unpatched, critical remote code execution flaw in LMCache, a popular add-on that speeds up vLLM inference servers; Lumen’s Black Lotus Labs exposed PoeLLM, a botnet that has hijacked more than 3,400 servers, led by internet-facing LiteLLM gateways, to mine cryptocurrency; contestants at Pwn2Own Ireland broke OpenAI Codex, LiteLLM and Oracle’s AI database for $135,000; and a poisoned Tensorlake npm package went hunting for the config files of AI coding tools. None of these needed a clever prompt. They needed an open port, a missing password or a trusted build pipeline. Below: what happened in each case, why the AI stack is so exposed, and a four-hour runbook for checking your own.
Key takeaways
- A 9.8 with no patch. CVE-2026-105192 lets anyone who can reach LMCache’s multiprocess server run code on it with a single message. Versions 0.3.9 through 0.5.5 are affected, and no fixed release exists yet.
- AI servers are being farmed. PoeLLM has compromised more than 3,400 servers since April, mostly exposed LiteLLM and Gotenberg deployments, and turns some victims into scanners that find the next ones.
- AI is now a Pwn2Own category. On day one in Ireland, researchers earned $40,000 each for exploits against OpenAI Codex, LiteLLM and Oracle Autonomous AI Database.
- Developer machines are the side door. The malicious Tensorlake 0.5.144 package stole cloud keys and AI tool configs, then planted files that re-run it when a project is opened in Claude Code or VS Code.
- The fix is old-fashioned. Inventory, no public exposure, authentication on every internal service, least-privilege containers and pinned dependencies stop most of what happened this week.
1. Four incidents, one pattern
Taken one at a time, each of this week’s stories looks like routine security news. Taken together, they show attackers working systematically through the layers that sit between a business and its AI models: the developer laptop, the gateway that routes requests to model providers, the inference server and its cache, and the databases and tools that agents call.
| Incident | Layer hit | What an attacker gets | Status on 8 Oct |
|---|---|---|---|
| LMCache CVE-2026-105192 (JFrog) | Inference cache next to vLLM | Code execution on the cache server, as root in official containers | Disclosed 7 Oct; no fixed version |
| PoeLLM / “Canto Incognito” (Lumen) | AI gateways (LiteLLM), document converters (Gotenberg), Gitea, Ivanti Sentry | Compute for cryptomining, a foothold, and a scanner for the next victim | Active since April; 3,400+ servers hit |
| Pwn2Own Ireland 2026, AI categories (ZDI) | Coding agents, AI gateways, AI databases | Proof that zero-days exist in OpenAI Codex, LiteLLM and Oracle Autonomous AI Database | Vendors have 90 days to patch |
| Tensorlake npm 0.5.144 (Socket, StepSecurity) | Developer machines and CI | Cloud, GitHub and npm tokens, AI tool configs, persistence | Version removed from npm; rotate secrets |
We saw the opening move a week ago in The AI Control Plane Is the New Crown Jewel, when GitLab patched a 9.9 flaw in its self-hosted AI Gateway. This week shows that was not a one-off.
2. LMCache: one message, full control, no patch
LMCache is open-source software that stores and reuses the intermediate results a large language model computes, so that repeated or overlapping prompts are answered faster and more cheaply. It is commonly paired with vLLM, one of the most widely used engines for running open-weight models on your own hardware. On 7 October, JFrog researcher Yuval Moravchick disclosed CVE-2026-105192, which JFrog scores 9.8 out of 10.
The bug lives in LMCache’s multiprocess mode, where the cache runs as a standalone server and model workers talk to it over a ZeroMQ socket. That socket has no authentication. One of its message types is decoded with Python’s pickle format, which can carry instructions that run during decoding, and the server decodes the message before it checks what kind of message it is. A single crafted message is enough to run the sender’s code. On the project’s official container images the process runs as root, according to JFrog.
There is some good news. By default the server listens only on the local machine, and a copy of LMCache embedded inside a single vLLM process opens no port at all. The risk appears when teams scale out: multi-node deployments sometimes bind the server to a network address, and the project’s own example Kubernetes DaemonSet for version 0.5.5 listens on every interface. Affected versions run from 0.3.9, released in October 2025, through 0.5.5, plus the 0.5.6 release candidates. As of 8 October there is no fixed version and no LMCache security advisory.
Two details make this worse. A day earlier, one GitHub user filed six more LMCache reports alleging cross-tenant data access and unauthenticated command execution; they are unconfirmed, but suggest more is coming. And the root cause, pickle over an unauthenticated socket, is the same design mistake researchers found across several AI inference frameworks in November 2025 under the “ShadowMQ” name.
What to do: keep the multiprocess server off routable addresses, restrict the port to the local host or a trusted cluster network, and remember JFrog’s warning that firewalls only shrink the problem: any host that can still connect can still run code. JFrog also offers no way to tell whether a server has already been hit, so treat any exposed instance as suspect.
3. PoeLLM: a botnet that farms AI servers
While JFrog was disclosing a theoretical risk, Lumen’s Black Lotus Labs was documenting a real campaign. Lumen calls the activity “Canto Incognito” and the malware PoeLLM. It has been running since April 2026, has compromised more than 3,400 servers, and peaked in mid-June at about 2,200 infected servers, with nearly 800 active on a given day. Victims are concentrated in the United States and Western Europe.
The targets are telling. PoeLLM mainly goes after internet-facing enterprise deployments of LiteLLM, an open-source gateway that gives applications one interface to many model providers and often holds the API keys for all of them, and Gotenberg, a document-conversion service common in AI document pipelines. It also hits Gitea code servers and Ivanti Sentry appliances. It breaks in through known, already-published vulnerabilities, installs the XMRig and Iron cryptominers connected to the Kryptex mining service, and then turns some victims into scanners that hunt for the next exposed instance.
The command-and-control trick is almost whimsical: the server address is hidden inside a poem posted to a GitHub repository, and the operators change a few words to point the botnet somewhere new. Lumen attributes the campaign, with moderate confidence, to an Italian-speaking actor. Its summary is the line every AI platform team should pin to the wall: “AI infrastructure is becoming an attractive target.”
Mining is the visible payload, but it is not the real cost. An AI gateway typically stores the keys for every model provider a company uses, plus logs of prompts and responses. A miner tells you the attacker had code execution; it does not tell you what else they took. If you find one, treat it as a breach, rotate every key the server held and check provider bills for unexplained usage. Our Shadow AI in 2026 guide explains why many of these gateways were stood up by enthusiastic teams without security review in the first place.
4. Pwn2Own makes AI a standing target
Pwn2Own, run by Trend Micro’s Zero Day Initiative (ZDI), pays researchers to break widely used products under controlled conditions and then gives vendors 90 days to fix the bugs before details are published. This year’s Ireland edition added dedicated categories for AI infrastructure and AI coding apps alongside phones, printers and smart-home kit. Day one, on 6 October, produced 32 unique zero-days and $388,500 in awards across 21 entries.
Ikotas Labs took down the cloud version of OpenAI’s Codex coding agent with a single argument-injection bug. Taisic Yun of Xint chained an input-validation weakness with code injection in LiteLLM to get a reverse shell, and a second team, Out of Bounds, used four LiteLLM bugs, two already known to the maintainers, for a smaller award. VinSOC chained five zero-days against Oracle Autonomous AI Database. Only Chroma, a vector database, survived its attempt, and only because the clock ran out.
These are contest exploits, not attacks on real users, and patches should follow within 90 days. The lesson is about maturity. Phones and browsers have had a decade of this kind of scrutiny; AI gateways and coding agents are getting it for the first time, and the first pass is finding a lot. For context on how AI-assisted research is also flooding maintainers with bug reports, see AI Broke the Bug Bounty.
5. The side door: a poisoned SDK that hunts AI tool configs
The fourth story arrived this morning. Version 0.5.144 of tensorlake, the npm package for Tensorlake’s TypeScript SDK, was published with a credential-stealing worm from the “Shai-Hulud” family. According to StepSecurity, the first rogue commit landed on the project’s main branch at 01:20 UTC on 7 October under a maintainer’s name, and the project’s own release workflow then published the poisoned version. That is the uncomfortable part: it came from the official source, through the official pipeline.
Socket’s analysis shows the payload runs from a preinstall hook and collects npm, GitHub, AWS, Kubernetes and HashiCorp Vault credentials, SSH keys, .env files, crypto wallets and, notably, the configuration and MCP files of Claude, Cursor, Kiro, Windsurf and Zed. AI coding tools now hold some of the most valuable tokens on a developer’s laptop. The worm then republishes itself into other packages the victim can publish, and writes .claude/settings.json and .vscode/tasks.json files into reachable repositories so that it runs again when someone opens the project in Claude Code or VS Code. A watchdog checks whether the stolen GitHub token still works and, if it is revoked, runs an attacker-supplied routine that earlier waves used destructively.
We covered how AI coding agents can be steered into leaking secrets in Encrypted Prompts Beat the Guardrails. Tensorlake shows the reverse route: skip the agent and steal its keys directly.
6. Why AI plumbing is so exposed
None of these flaws is exotic. Unauthenticated internal sockets, unsafe deserialisation, argument injection and poisoned packages are problems the industry has known how to prevent for years. Four habits explain why they keep turning up in AI software:
- Research code in production. Many inference tools started as fast-moving research projects where speed mattered more than hardening. Pickle is convenient; authentication is extra work.
- “Internal” assumptions that do not survive scaling. A component designed to talk to its neighbour on localhost gets a network address the day a team spreads it across a GPU cluster.
- Keys concentrated in one place. Gateways and coding tools hold the credentials for every model provider and cloud account, which makes them worth far more than the GPU time a miner steals.
- No owner. Gateways, vector databases and MCP servers are often installed by product or data teams, not IT, so they miss the patch cycle. OX Security’s scan of 15,465 public MCP servers this week found 2.3% of hostnames no longer resolve, and six sit on expired domains anyone could buy for a few dollars a year, inheriting the agents still configured to call them.
Speed makes this urgent. As we covered in The 24-Hour Window, attackers now move from disclosure to exploitation in about a day. LMCache’s flaw is fully described in public, and PoeLLM already has the scanning machinery to look for new targets.
7. The four-hour AI infrastructure runbook
You do not need a programme to respond to this week. You need one focused afternoon. Split it into four one-hour blocks and do them in order; each block ends with something you can show your manager.
| Hour | Goal | Do this | Done when |
|---|---|---|---|
| Hour 1 | Find it | Search cloud accounts, Kubernetes clusters and container registries for LMCache, vLLM, LiteLLM, Gotenberg, Chroma and self-hosted MCP servers. Ask product and data teams directly; many of these were never ticketed. | A list with owner, version and network exposure for each instance |
| Hour 2 | Close the doors | Bind LMCache multiprocess servers to localhost or a private cluster network. Take LiteLLM and Gotenberg off the public internet or put them behind a VPN or authenticated proxy. Check the example DaemonSet has not been copied unchanged. | No AI component answers from the internet without authentication |
| Hour 3 | Look for intruders | Check AI hosts for unexpected CPU spikes, XMRig or Iron processes, outbound mining connections and new SSH keys. Search lockfiles and CI logs for tensorlake 0.5.144, and repositories for unexpected .claude/settings.json or .vscode/tasks.json files. | Each host marked clean, suspect or isolated |
| Hour 4 | Contain and harden | For anything suspect: isolate first, then rotate model-provider keys, cloud credentials and GitHub tokens. Run AI containers as non-root, pin package versions, and set billing alerts on every model-provider account. | Keys rotated, alerts on, a date set to recheck for the LMCache patch |
Two follow-ups belong on next week’s list. First, give every AI component an owner and put it in the normal patch cycle; the 90-day Pwn2Own clock means LiteLLM, Codex and Oracle fixes are coming, and someone needs to apply them. Second, add AI tool configs to your secrets policy. Our AI Agent Security in 2026 guide covers how to assign that ownership.
8. What to watch next
An LMCache fix. Watch the project’s GitHub releases and security advisories for a patched version, and for responses to the six unverified reports. Until then, configuration is the only defence.
Whether PoeLLM adds LMCache. Botnets that already scan for exposed AI services tend to add fresh, public exploits quickly. If Lumen or others report LMCache exploitation, move it to the top of the queue. We have seen the same pattern of AI-assisted attackers outpacing patch cycles in AI Agents Forced a Security Rewrite.
Frequently asked questions
What is CVE-2026-105192?
It is a critical flaw in LMCache’s multiprocess mode, disclosed by JFrog on 7 October 2026 and scored 9.8. The server’s ZeroMQ socket has no authentication and decodes pickled data before checking the message type, so anyone who can reach it can run code. Versions 0.3.9 through 0.5.5 and the 0.5.6 release candidates are affected, and no fix exists yet.
Am I affected if I use vLLM?
Only if you run LMCache in multiprocess mode and the server is reachable from other hosts. The server listens on localhost by default, and LMCache embedded in a single vLLM process opens no port. Multi-node deployments, including the project’s example Kubernetes DaemonSet, are the main risk.
What is the PoeLLM botnet?
PoeLLM is malware documented by Lumen’s Black Lotus Labs that has compromised more than 3,400 servers since April 2026, mainly exposed LiteLLM and Gotenberg deployments plus Gitea and Ivanti Sentry. It installs cryptominers and uses infected servers to scan for new victims.
Is LiteLLM safe to use?
It can be, if it is patched, kept off the public internet and protected by authentication. PoeLLM exploits known vulnerabilities in internet-facing deployments, and Pwn2Own showed new zero-days that should be fixed within ZDI’s 90-day window. Treat it as a high-value system, because it usually holds every model-provider key you use.
What should I do if I installed tensorlake 0.5.144?
Remove it, isolate the affected machine and plan credential rotation before revoking tokens, because the worm includes a watchdog that can trigger a harmful routine when its stolen GitHub token stops working. Then rotate npm, GitHub, cloud, Vault, Kubernetes and SSH credentials, and check repositories for unexpected .claude/settings.json, .vscode/tasks.json and GitHub Actions files.
Why are attackers targeting AI infrastructure now?
AI servers have powerful GPUs and CPUs worth stealing, gateways store valuable API keys, and much of the software is young and often deployed without security review. Combined with faster exploitation, that makes exposed AI components an easy, profitable target.
Sources
- The Hacker News: Unpatched Critical LMCache Flaw Lets Unauthenticated Attackers Run Code Remotely (7 Oct 2026)
- The Hacker News: PoeLLM Malware Infects 3,400+ Servers to Expand Crypto Mining Botnet (7 Oct 2026)
- Cyber Security News: 32 zero-days exploited on day one of Pwn2Own Ireland 2026 (7 Oct 2026)
- BleepingComputer: Hackers exploit 32 zero-days on first day of Pwn2Own Ireland (6 Oct 2026)
- The Hacker News: Tensorlake npm Package Compromised to Deliver Shai-Hulud Credential-Stealing Worm (8 Oct 2026)
- The Hacker News / OX Security: What We Found Inside 15,465 Public MCP Servers (6 Oct 2026)
