Skip to content

Mon - Fri: 10.00 - 5.00

[email protected]

Delana Technologies

Delana Technologies

Delana Technologies delivers expert cybersecurity, cloud, and AI-driven IT strategy solutions. Transform your enterprise securely and intelligently.

  • Home
  • Contact Us
  • About Us
  • Case Studies
  • Workflow Automation & Systems Integration
  • AI Consulting & Agentic AI Solutions
  • Home
  • Contact Us
  • About Us
  • Case Studies
  • Workflow Automation & Systems Integration
  • AI Consulting & Agentic AI Solutions

Mon - Fri: 10.00 - 5.00

[email protected]

Your AI Stack Is the New Target: An Unpatched 9.8 LMCache Flaw, a 3,400-Server LLM Botnet and Pwn2Own’s AI Hacks (AI Trends, 8 October 2026)

  1. Home   »  
  2. Your AI Stack Is the New Target: An Unpatched 9.8 LMCache Flaw, a 3,400-Server LLM Botnet and Pwn2Own’s AI Hacks (AI Trends, 8 October 2026)

Your AI Stack Is the New Target: An Unpatched 9.8 LMCache Flaw, a 3,400-Server LLM Botnet and Pwn2Own’s AI Hacks (AI Trends, 8 October 2026)

October 8, 2026October 8, 2026 admincybersecurityTagged AI gateway, AI infrastructure security, AI security, AI trends, LiteLLM, LMCache, MCP servers, open source security, Pwn2Own, secrets management, software supply chain, vulnerability management

What’s trending in AI on 8 October 2026: For two years the AI security conversation has been about what models say and what agents do. This week the attackers went for the plumbing underneath. In 48 hours, researchers disclosed an unpatched, critical remote code execution flaw in LMCache, a popular add-on that speeds up vLLM inference servers; Lumen’s Black Lotus Labs exposed PoeLLM, a botnet that has hijacked more than 3,400 servers, led by internet-facing LiteLLM gateways, to mine cryptocurrency; contestants at Pwn2Own Ireland broke OpenAI Codex, LiteLLM and Oracle’s AI database for $135,000; and a poisoned Tensorlake npm package went hunting for the config files of AI coding tools. None of these needed a clever prompt. They needed an open port, a missing password or a trusted build pipeline. Below: what happened in each case, why the AI stack is so exposed, and a four-hour runbook for checking your own.

Key takeaways

  • A 9.8 with no patch. CVE-2026-105192 lets anyone who can reach LMCache’s multiprocess server run code on it with a single message. Versions 0.3.9 through 0.5.5 are affected, and no fixed release exists yet.
  • AI servers are being farmed. PoeLLM has compromised more than 3,400 servers since April, mostly exposed LiteLLM and Gotenberg deployments, and turns some victims into scanners that find the next ones.
  • AI is now a Pwn2Own category. On day one in Ireland, researchers earned $40,000 each for exploits against OpenAI Codex, LiteLLM and Oracle Autonomous AI Database.
  • Developer machines are the side door. The malicious Tensorlake 0.5.144 package stole cloud keys and AI tool configs, then planted files that re-run it when a project is opened in Claude Code or VS Code.
  • The fix is old-fashioned. Inventory, no public exposure, authentication on every internal service, least-privilege containers and pinned dependencies stop most of what happened this week.
Your AI stack is the new targetTitle card. Headline: Your AI stack is the new target. Subhead: attackers skipped the model and went for the plumbing. Three tags: LMCache 9.8 flaw, unpatched; PoeLLM botnet, 3,400 plus servers; Pwn2Own AI hacks, 135,000 dollars. Illustration of a stacked server tower with four layers labelled dev tools, gateway, inference, data, with orange warning markers on each layer. DEV TOOLS! AI GATEWAY! INFERENCE + CACHE! DATA + TOOLS! AI TRENDS · 8 OCTOBER 2026 Your AI stack is the new target Attackers skipped the model this week and went straight for the plumbing. LMCache 9.8 flaw · no patch yet PoeLLM botnet · 3,400+ servers Pwn2Own AI hacks · $135,000 Sources: JFrog via The Hacker News, Lumen Black Lotus Labs, ZDI via Cyber Security News (6–8 Oct 2026)delana.co
Four layers, four incidents in 48 hours. None of them involved tricking a model.

1. Four incidents, one pattern

Taken one at a time, each of this week’s stories looks like routine security news. Taken together, they show attackers working systematically through the layers that sit between a business and its AI models: the developer laptop, the gateway that routes requests to model providers, the inference server and its cache, and the databases and tools that agents call.

IncidentLayer hitWhat an attacker getsStatus on 8 Oct
LMCache CVE-2026-105192 (JFrog)Inference cache next to vLLMCode execution on the cache server, as root in official containersDisclosed 7 Oct; no fixed version
PoeLLM / “Canto Incognito” (Lumen)AI gateways (LiteLLM), document converters (Gotenberg), Gitea, Ivanti SentryCompute for cryptomining, a foothold, and a scanner for the next victimActive since April; 3,400+ servers hit
Pwn2Own Ireland 2026, AI categories (ZDI)Coding agents, AI gateways, AI databasesProof that zero-days exist in OpenAI Codex, LiteLLM and Oracle Autonomous AI DatabaseVendors have 90 days to patch
Tensorlake npm 0.5.144 (Socket, StepSecurity)Developer machines and CICloud, GitHub and npm tokens, AI tool configs, persistenceVersion removed from npm; rotate secrets
Compiled from The Hacker News, Cyber Security News and BleepingComputer reporting, 6–8 October 2026.

We saw the opening move a week ago in The AI Control Plane Is the New Crown Jewel, when GitLab patched a 9.9 flaw in its self-hosted AI Gateway. This week shows that was not a one-off.

2. LMCache: one message, full control, no patch

LMCache is open-source software that stores and reuses the intermediate results a large language model computes, so that repeated or overlapping prompts are answered faster and more cheaply. It is commonly paired with vLLM, one of the most widely used engines for running open-weight models on your own hardware. On 7 October, JFrog researcher Yuval Moravchick disclosed CVE-2026-105192, which JFrog scores 9.8 out of 10.

The bug lives in LMCache’s multiprocess mode, where the cache runs as a standalone server and model workers talk to it over a ZeroMQ socket. That socket has no authentication. One of its message types is decoded with Python’s pickle format, which can carry instructions that run during decoding, and the server decodes the message before it checks what kind of message it is. A single crafted message is enough to run the sender’s code. On the project’s official container images the process runs as root, according to JFrog.

How the LMCache flaw worksFour-step attack path. Step 1: the LMCache multiprocess server is bound to a routable address, for example via the example Kubernetes DaemonSet that listens on all interfaces. Step 2: an attacker sends one ZeroMQ message with no authentication. Step 3: the server unpickles the message arguments before checking the message type. Step 4: attacker code runs as the LMCache process, which is root in official container images. Below, a safe configuration: bound to localhost only, which other hosts cannot reach. Affected versions 0.3.9 to 0.5.5 and 0.5.6 release candidates; no fix as of 8 October 2026. One message to root: the LMCache flaw CVE-2026-105192 · severity 9.8 · versions 0.3.9 to 0.5.5 and 0.5.6 release candidates 1 · EXPOSEDServer bound to aroutable address(e.g. example DaemonSet) 2 · NO AUTHAttacker sends oneZeroMQ message 3 · UNPICKLE FIRSTArguments decodedbefore type check 4 · CODE RUNSAs the LMCache user:root in official images Safe by default, unsafe when scaled out The server listens on localhost unless an operator changes it. LMCache embedded inside a single vLLM process opens no port. Multi-node clusters are where the risk appears. Source: JFrog disclosure via The Hacker News (7 Oct 2026)delana.co
The weak link is the assumption that “internal” traffic is trusted traffic.

There is some good news. By default the server listens only on the local machine, and a copy of LMCache embedded inside a single vLLM process opens no port at all. The risk appears when teams scale out: multi-node deployments sometimes bind the server to a network address, and the project’s own example Kubernetes DaemonSet for version 0.5.5 listens on every interface. Affected versions run from 0.3.9, released in October 2025, through 0.5.5, plus the 0.5.6 release candidates. As of 8 October there is no fixed version and no LMCache security advisory.

Two details make this worse. A day earlier, one GitHub user filed six more LMCache reports alleging cross-tenant data access and unauthenticated command execution; they are unconfirmed, but suggest more is coming. And the root cause, pickle over an unauthenticated socket, is the same design mistake researchers found across several AI inference frameworks in November 2025 under the “ShadowMQ” name.

What to do: keep the multiprocess server off routable addresses, restrict the port to the local host or a trusted cluster network, and remember JFrog’s warning that firewalls only shrink the problem: any host that can still connect can still run code. JFrog also offers no way to tell whether a server has already been hit, so treat any exposed instance as suspect.

3. PoeLLM: a botnet that farms AI servers

While JFrog was disclosing a theoretical risk, Lumen’s Black Lotus Labs was documenting a real campaign. Lumen calls the activity “Canto Incognito” and the malware PoeLLM. It has been running since April 2026, has compromised more than 3,400 servers, and peaked in mid-June at about 2,200 infected servers, with nearly 800 active on a given day. Victims are concentrated in the United States and Western Europe.

The targets are telling. PoeLLM mainly goes after internet-facing enterprise deployments of LiteLLM, an open-source gateway that gives applications one interface to many model providers and often holds the API keys for all of them, and Gotenberg, a document-conversion service common in AI document pipelines. It also hits Gitea code servers and Ivanti Sentry appliances. It breaks in through known, already-published vulnerabilities, installs the XMRig and Iron cryptominers connected to the Kryptex mining service, and then turns some victims into scanners that hunt for the next exposed instance.

PoeLLM by the numbersBar chart of PoeLLM scale. Nearly 800 servers active per day at peak; about 2,200 affected servers at the mid-June peak; more than 3,400 servers identified in total since April 2026. To the right, a loop showing the botnet cycle: exploit a known flaw in an exposed LiteLLM, Gotenberg, Gitea or Ivanti Sentry server; install XMRig and Iron miners; turn the host into a scanner; find the next victim. PoeLLM by the numbers Servers compromised by the “Canto Incognito” campaign, April to October 2026 Active per day at peak~800 Affected at mid-June peak~2,200 Total identified so far3,400+ Exploit knownflaw Installminers Become ascanner Find nextvictim Source: Lumen Black Lotus Labs via The Hacker News (7 Oct 2026)delana.co
A self-feeding loop: every exposed AI server it captures helps it find the next one.

The command-and-control trick is almost whimsical: the server address is hidden inside a poem posted to a GitHub repository, and the operators change a few words to point the botnet somewhere new. Lumen attributes the campaign, with moderate confidence, to an Italian-speaking actor. Its summary is the line every AI platform team should pin to the wall: “AI infrastructure is becoming an attractive target.”

Mining is the visible payload, but it is not the real cost. An AI gateway typically stores the keys for every model provider a company uses, plus logs of prompts and responses. A miner tells you the attacker had code execution; it does not tell you what else they took. If you find one, treat it as a breach, rotate every key the server held and check provider bills for unexplained usage. Our Shadow AI in 2026 guide explains why many of these gateways were stood up by enthusiastic teams without security review in the first place.

4. Pwn2Own makes AI a standing target

Pwn2Own, run by Trend Micro’s Zero Day Initiative (ZDI), pays researchers to break widely used products under controlled conditions and then gives vendors 90 days to fix the bugs before details are published. This year’s Ireland edition added dedicated categories for AI infrastructure and AI coding apps alongside phones, printers and smart-home kit. Day one, on 6 October, produced 32 unique zero-days and $388,500 in awards across 21 entries.

Pwn2Own Ireland 2026 day one: AI targetsHorizontal bar chart of day-one awards in the AI categories. OpenAI Codex, single argument-injection bug, Ikotas Labs: 40,000 dollars. LiteLLM, input validation plus code injection to a reverse shell, Taisic Yun of Xint: 40,000 dollars. Oracle Autonomous AI Database, five chained flaws, VinSOC: 40,000 dollars. LiteLLM, four bugs including two already known, Out of Bounds: 15,000 dollars. Chroma, attempt ran out of time: zero. AI category total 135,000 dollars of 388,500 dollars on day one. Pwn2Own Ireland, day one: the AI targets Awards for AI infrastructure and AI coding app exploits, 6 October 2026 OpenAI Codex1 argument-injection bug · Ikotas Labs$40,000 LiteLLMValidation + code injection · Xint$40,000 Oracle Autonomous AI DB5 chained flaws · VinSOC$40,000 LiteLLM (second team)4 bugs, 2 already known · Out of Bounds$15,000 ChromaRan out of time · VinSOC$0 AI total: $135,000 Source: ZDI day-one results via Cyber Security News and BleepingComputer (6–7 Oct 2026)delana.co
LiteLLM fell twice on day one, the same software PoeLLM has been exploiting in the wild since April.

Ikotas Labs took down the cloud version of OpenAI’s Codex coding agent with a single argument-injection bug. Taisic Yun of Xint chained an input-validation weakness with code injection in LiteLLM to get a reverse shell, and a second team, Out of Bounds, used four LiteLLM bugs, two already known to the maintainers, for a smaller award. VinSOC chained five zero-days against Oracle Autonomous AI Database. Only Chroma, a vector database, survived its attempt, and only because the clock ran out.

These are contest exploits, not attacks on real users, and patches should follow within 90 days. The lesson is about maturity. Phones and browsers have had a decade of this kind of scrutiny; AI gateways and coding agents are getting it for the first time, and the first pass is finding a lot. For context on how AI-assisted research is also flooding maintainers with bug reports, see AI Broke the Bug Bounty.

5. The side door: a poisoned SDK that hunts AI tool configs

The fourth story arrived this morning. Version 0.5.144 of tensorlake, the npm package for Tensorlake’s TypeScript SDK, was published with a credential-stealing worm from the “Shai-Hulud” family. According to StepSecurity, the first rogue commit landed on the project’s main branch at 01:20 UTC on 7 October under a maintainer’s name, and the project’s own release workflow then published the poisoned version. That is the uncomfortable part: it came from the official source, through the official pipeline.

Socket’s analysis shows the payload runs from a preinstall hook and collects npm, GitHub, AWS, Kubernetes and HashiCorp Vault credentials, SSH keys, .env files, crypto wallets and, notably, the configuration and MCP files of Claude, Cursor, Kiro, Windsurf and Zed. AI coding tools now hold some of the most valuable tokens on a developer’s laptop. The worm then republishes itself into other packages the victim can publish, and writes .claude/settings.json and .vscode/tasks.json files into reachable repositories so that it runs again when someone opens the project in Claude Code or VS Code. A watchdog checks whether the stolen GitHub token still works and, if it is revoked, runs an attacker-supplied routine that earlier waves used destructively.

We covered how AI coding agents can be steered into leaking secrets in Encrypted Prompts Beat the Guardrails. Tensorlake shows the reverse route: skip the agent and steal its keys directly.

6. Why AI plumbing is so exposed

None of these flaws is exotic. Unauthenticated internal sockets, unsafe deserialisation, argument injection and poisoned packages are problems the industry has known how to prevent for years. Four habits explain why they keep turning up in AI software:

  • Research code in production. Many inference tools started as fast-moving research projects where speed mattered more than hardening. Pickle is convenient; authentication is extra work.
  • “Internal” assumptions that do not survive scaling. A component designed to talk to its neighbour on localhost gets a network address the day a team spreads it across a GPU cluster.
  • Keys concentrated in one place. Gateways and coding tools hold the credentials for every model provider and cloud account, which makes them worth far more than the GPU time a miner steals.
  • No owner. Gateways, vector databases and MCP servers are often installed by product or data teams, not IT, so they miss the patch cycle. OX Security’s scan of 15,465 public MCP servers this week found 2.3% of hostnames no longer resolve, and six sit on expired domains anyone could buy for a few dollars a year, inheriting the agents still configured to call them.
Map of the AI stack and this week’s incidentsFive-layer stack diagram. Layer 1, developer machines and CI: Tensorlake npm worm and Codex argument injection; control: pinned dependencies and short-lived tokens. Layer 2, AI gateway: PoeLLM botnet and Pwn2Own LiteLLM exploits; control: no public exposure, patching, key vault. Layer 3, inference and cache: LMCache CVE-2026-105192; control: localhost binding, authentication, non-root containers. Layer 4, data and vector stores: Oracle Autonomous AI Database and Chroma at Pwn2Own; control: network isolation and least privilege. Layer 5, agent tools and MCP servers: 2.3 percent of public MCP hostnames dangling; control: approved server list and domain monitoring. Where this week’s incidents landed Each layer of a typical AI stack, what hit it, and the control that matters most LAYERTHIS WEEKKEY CONTROL Developer machines + CITensorlake worm · Codex bugPinned deps, short-lived tokens AI gateway (LiteLLM)PoeLLM botnet · 2 Pwn2Own hacksNever public, patch, vault keys Inference + cacheLMCache 9.8 · unpatchedLocalhost only, non-root Data + vector storesOracle AI DB chain · Chroma tryIsolate, least privilege Agent tools + MCP2.3% of MCP hosts danglingApproved list, watch domains Sources: JFrog, Lumen, ZDI, Socket, StepSecurity, OX Security (6–8 Oct 2026); controls are Delana’s recommendationsdelana.co
Every layer was touched this week. The inference cache is the only one with no patch available.

Speed makes this urgent. As we covered in The 24-Hour Window, attackers now move from disclosure to exploitation in about a day. LMCache’s flaw is fully described in public, and PoeLLM already has the scanning machinery to look for new targets.

7. The four-hour AI infrastructure runbook

You do not need a programme to respond to this week. You need one focused afternoon. Split it into four one-hour blocks and do them in order; each block ends with something you can show your manager.

HourGoalDo thisDone when
Hour 1Find itSearch cloud accounts, Kubernetes clusters and container registries for LMCache, vLLM, LiteLLM, Gotenberg, Chroma and self-hosted MCP servers. Ask product and data teams directly; many of these were never ticketed.A list with owner, version and network exposure for each instance
Hour 2Close the doorsBind LMCache multiprocess servers to localhost or a private cluster network. Take LiteLLM and Gotenberg off the public internet or put them behind a VPN or authenticated proxy. Check the example DaemonSet has not been copied unchanged.No AI component answers from the internet without authentication
Hour 3Look for intrudersCheck AI hosts for unexpected CPU spikes, XMRig or Iron processes, outbound mining connections and new SSH keys. Search lockfiles and CI logs for tensorlake 0.5.144, and repositories for unexpected .claude/settings.json or .vscode/tasks.json files.Each host marked clean, suspect or isolated
Hour 4Contain and hardenFor anything suspect: isolate first, then rotate model-provider keys, cloud credentials and GitHub tokens. Run AI containers as non-root, pin package versions, and set billing alerts on every model-provider account.Keys rotated, alerts on, a date set to recheck for the LMCache patch
Delana’s runbook, based on guidance from JFrog, Lumen, Socket and StepSecurity. Adapt the inventory to the AI tools you actually run.

Two follow-ups belong on next week’s list. First, give every AI component an owner and put it in the normal patch cycle; the 90-day Pwn2Own clock means LiteLLM, Codex and Oracle fixes are coming, and someone needs to apply them. Second, add AI tool configs to your secrets policy. Our AI Agent Security in 2026 guide covers how to assign that ownership.

8. What to watch next

An LMCache fix. Watch the project’s GitHub releases and security advisories for a patched version, and for responses to the six unverified reports. Until then, configuration is the only defence.

Whether PoeLLM adds LMCache. Botnets that already scan for exposed AI services tend to add fresh, public exploits quickly. If Lumen or others report LMCache exploitation, move it to the top of the queue. We have seen the same pattern of AI-assisted attackers outpacing patch cycles in AI Agents Forced a Security Rewrite.

Frequently asked questions

What is CVE-2026-105192?

It is a critical flaw in LMCache’s multiprocess mode, disclosed by JFrog on 7 October 2026 and scored 9.8. The server’s ZeroMQ socket has no authentication and decodes pickled data before checking the message type, so anyone who can reach it can run code. Versions 0.3.9 through 0.5.5 and the 0.5.6 release candidates are affected, and no fix exists yet.

Am I affected if I use vLLM?

Only if you run LMCache in multiprocess mode and the server is reachable from other hosts. The server listens on localhost by default, and LMCache embedded in a single vLLM process opens no port. Multi-node deployments, including the project’s example Kubernetes DaemonSet, are the main risk.

What is the PoeLLM botnet?

PoeLLM is malware documented by Lumen’s Black Lotus Labs that has compromised more than 3,400 servers since April 2026, mainly exposed LiteLLM and Gotenberg deployments plus Gitea and Ivanti Sentry. It installs cryptominers and uses infected servers to scan for new victims.

Is LiteLLM safe to use?

It can be, if it is patched, kept off the public internet and protected by authentication. PoeLLM exploits known vulnerabilities in internet-facing deployments, and Pwn2Own showed new zero-days that should be fixed within ZDI’s 90-day window. Treat it as a high-value system, because it usually holds every model-provider key you use.

What should I do if I installed tensorlake 0.5.144?

Remove it, isolate the affected machine and plan credential rotation before revoking tokens, because the worm includes a watchdog that can trigger a harmful routine when its stolen GitHub token stops working. Then rotate npm, GitHub, cloud, Vault, Kubernetes and SSH credentials, and check repositories for unexpected .claude/settings.json, .vscode/tasks.json and GitHub Actions files.

Why are attackers targeting AI infrastructure now?

AI servers have powerful GPUs and CPUs worth stealing, gateways store valuable API keys, and much of the software is young and often deployed without security review. Combined with faster exploitation, that makes exposed AI components an easy, profitable target.


Sources

  • The Hacker News: Unpatched Critical LMCache Flaw Lets Unauthenticated Attackers Run Code Remotely (7 Oct 2026)
  • The Hacker News: PoeLLM Malware Infects 3,400+ Servers to Expand Crypto Mining Botnet (7 Oct 2026)
  • Cyber Security News: 32 zero-days exploited on day one of Pwn2Own Ireland 2026 (7 Oct 2026)
  • BleepingComputer: Hackers exploit 32 zero-days on first day of Pwn2Own Ireland (6 Oct 2026)
  • The Hacker News: Tensorlake npm Package Compromised to Deliver Shai-Hulud Credential-Stealing Worm (8 Oct 2026)
  • The Hacker News / OX Security: What We Found Inside 15,465 Public MCP Servers (6 Oct 2026)

Post navigation

Previous: AI Agents Get a Front Door: Meta and Sierra’s Personal Agent Protocol Wants Bots to Stop Pretending to Be You (AI Trends, 7 October 2026)

Florida Service Location

  • Cybersecurity, AI Consulting & IT Services in West Palm Beach, Florida
  • Cybersecurity, AI Consulting & IT Services in Sarasota, Florida
  • Cybersecurity, AI Consulting & IT Services in Port St. Lucie, Florida
  • Cybersecurity, AI Consulting & IT Services in Pembroke Pines, Florida
  • Cybersecurity, AI Consulting & IT Services in Naples, Florida
  • Cybersecurity, AI Consulting & IT Services in Miramar, Florida
  • Cybersecurity, AI Consulting & IT Services in Miami, Florida
  • Cybersecurity, AI Consulting & IT Services in Hollywood, Florida
  • Cybersecurity, AI Consulting & IT Services in Hialeah, Florida
  • Cybersecurity, AI Consulting & IT Services in Fort Myers, Florida
  • Cybersecurity, AI Consulting & IT Services in Fort Lauderdale, Florida
  • Cybersecurity, AI Consulting & IT Services in Cape Coral, Florida
  • Cybersecurity, AI Consulting & IT Services in Boca Raton, Florida
  • Cybersecurity, AI Consulting & IT Services in Coral Springs, Florida

Technology Services

  • Cybersecurity Compliance & Regulatory Framework Services
  • Workflow Automation & Systems Integration
  • Cloud Modernization & Technology Innovation Services
  • Fractional CTO & Expert Technical Consultants
  • Data Analytics, BI & Modern Data Platforms
  • Cyber Litigation Support & Digital Forensics
  • Cybersecurity Solutions & Zero Trust Architecture
  • AI Consulting & Agentic AI Solutions
  • Case Studies
  • Home
  • Contact Us
  • Privacy Policy
  • Cybersecurity Compliance & Regulatory Framework Services
  • Workflow Automation & Systems Integration
  • Cloud Modernization & Technology Innovation Services
  • Fractional CTO & Expert Technical Consultants
  • Data Analytics, BI & Modern Data Platforms
  • Cyber Litigation Support & Digital Forensics
  • Cybersecurity Solutions & Zero Trust Architecture
  • AI Consulting & Agentic AI Solutions

© Copyright 2025 Delana Technologies LLC