vLLM (5 Oct CVE batch, CVE-2026-105752 to 105760): nine advisories, mostly single-request crashes of the shared EngineCore and unauthenticated memory exhaustion, plus a cross-tenant prefix-cache oracle — fixed in 0.30.0
Nine CVE records for vLLM, the open-source LLM inference and serving engine, were published on 5 October 2026 from its GitHub advisories; all are Medium or Low (top CVSS 3.1 score 6.5). The higher-impact ones let one request take down service for every user of a deployment: CVE-2026-105756 (GHSA-2823-qmq8-rwvj), where a malformed cache_salt kills EngineCore on LMCache-MP deployments; CVE-2026-105757 (GHSA-85xf-c7hm-whqw), where structured-output request errors escape the request boundary and terminate the shared engine; CVE-2026-105754, where the scale-out disaggregated multimodal transport trusts caller-supplied features; and CVE-2026-105753, a multimodal IPC cache desync (fixed in 0.28.0). Several are reachable without authentication: CVE-2026-105758 and CVE-2026-105760 let request-supplied video frame and fps values exhaust CPU and memory (Qwen2-VL/Qwen3-VL samplers via /tokenize, and the GLMGA backend), and CVE-2026-105759 lets arbitrary HTTP method tokens grow Prometheus label sets in the Rust frontend until memory runs out. Two weaken tenant isolation: CVE-2026-105752 drops cache_salt on Harmony tool continuations so a tenant can test whether another tenant's prefix was cached (prefix caching is on by default), and CVE-2026-105755 lets a reused X-Request-Id overwrite cached query embeddings on /score and /rerank. The advisories do not report exploitation. Primary: vLLM GitHub security advisories.
- Product
- vLLM (LLM inference and serving engine; OpenAI-compatible API, Rust frontend, multimodal and LMCache-MP paths)
- Versions
- CVE-2026-105752 and CVE-2026-105754 to CVE-2026-105760: before 0.30.0 (105758 from 0.24.0; 105760 from 0.23.0rc2). CVE-2026-105753: before 0.28.0. Fixed: 0.30.0 or later (current PyPI release is 0.31.0).
- CVSS
- highest (CVSS 3.1: CVE-2026-105753, 105754, 105756, 105757); others 3.1 to 5.9
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H - Exploited in Australia?
- unknown
- Patch to
- Upgrade vLLM to 0.30.0 or later. Put the serving API behind an authenticating gateway with request-size and rate limits, do not expose /tokenize or /metrics publicly, and on shared multi-tenant deployments enforce cache_salt per tenant at the gateway rather than trusting the client.
Primary: vLLM GHSA-85xf-c7hm-whqw — structured-output errors terminate shared EngineCore (CVE-2026-105757) · Vendor: vLLM — GitHub security advisories index · CVE: CVE-2026-105752, CVE-2026-105756, CVE-2026-105757, CVE-2026-105754, CVE-2026-105753, CVE-2026-105758, CVE-2026-105760, CVE-2026-105759, CVE-2026-105755 · vLLM GHSA-2823-qmq8-rwvj — cache_salt validation kills EngineCore on LMCache-MP (CVE-2026-105756)
