CVE-2026-105756: vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP deployments — uncaught downstream `ValueError` denial of service
Affected
- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the loose cachesalt validator and unguarded scheduling-path lookup reach.
Summary
vLLM's OpenAI-compatible request models (Completions, Chat Completions, Responses) accept a client-supplied cachesalt field and validate it only as "must be a non-empty string" — no character or length restrictions. On a deployment with the built-in LMCache-MP KV connector enabled, that value is stored verbatim on the request tracker and forwarded unguarded as a keyword argument into the scheduler's per-step cache lookup. The downstream LMCache library applies a stricter check in IPCCacheServerKey.postinit, which raises ValueError for any cachesalt containing @, /, \, or NUL (or longer than 128 characters).
Neither the LMCache-MP connector lookup call site nor Scheduler.schedule() wraps that call in a request-scoped try/except, so the ValueError propagates uncaught into EngineCore's top-level handler, which treats any uncaught exception as fatal and kills the whole engine process. A single publicly reachable request with, for example, cachesalt="/" therefore takes down the engine for all concurrent users. vLLM's boundary validator is looser than the downstream consumer's, and the gap is never converted into a request-scoped failure on the scheduling path.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
- The three checkcachesaltsupport validators require only a non-empty string — no character or length bound: vllm/entrypoints/openai/completion/protocol.py#L502-L508 (field at #L172), vllm/entrypoints/openai/chatcompletion/protocol.py#L913-L919 (field at #L425), and vllm/entrypoints/openai/responses/protocol.py#L459-L465 (field at #L235). The loose test itself is at completion #L503-L505, chatcompletion #L914-L916, responses #L460-L462. - Two further request models accept cachesalt with the same or weaker checking, and should be hardened at the same time: vllm/entrypoints/pooling/base/protocol.py#L74-L85 carries the identical non-empty-string-only validator (field at #L59), and the token-in-token-out scale-out GenerateRequest exposes the field with no cachesalt validator at all: vllm/entrypoints/scaleout/tokenintokenout/protocol.py#L110. - The LMCache-MP request tracker stores the salt verbatim: vllm/distributed/kvtransfer/kvconnector/v1/lmcachempconnector.py#L189 (LMCacheMPRequestTracker), assignment at #L223. - The lookup call forwards cachesalt with no surrounding try: vllm/distributed/kvtransfer/kvconnector/v1/lmcachempconnector.py#L733-L774 (getnumnewmatchedtokens → maybesubmitlookuprequest(...) at #L770). - The scheduler invokes the connector unguarded: vllm/v1/core/sched/scheduler.py#L739 (self.connector.getnumnewmatchedtokens(...), inside schedule() at #L396). - The generic top-level handler treats the exception as fatal: vllm/v1/engine/core.py#L1229-L1235 — the except Exception in runenginecore (#L1154) logs EngineCore encountered a fatal error., calls sendenginedead(), and re-raises. - Downstream strict validator (external LMCache library, not vLLM): IPCCacheServerKey.postinit in lmcache/v1/multiprocess/customtypes.py rejects @ / \ NUL and >128-char salts (SALTFORBIDDENCHARS = frozenset("@/\\\x00")) by raising ValueError.
The three OpenAI checkcachesaltsupport validators are identical; the completion one is representative — a non-empty-string test with no character or length bound:
python vllm/entrypoints/openai/completion/protocol.py Lines 503-509 if data.get("cachesalt") is not None and ( not isinstance(data["cachesalt"], str) or not data["cachesalt"] ): raise VLLMValidationError( "Parameter 'cachesalt' must be a non-empty string if provided.", parameter="cachesalt", )
On the scheduling path the connector is invoked with no surrounding try — a ValueError from the downstream salt check propagates straight out of schedule():
python vllm/v1/core/sched/scheduler.py Lines 736-742 # Get externally-cached tokens if using a KVConnector. if self.connector is not None: exttokens, loadkvasync = ( self.connector.getnumnewmatchedtokens( request, numnewlocalcomputedtokens ) )
runenginecore's generic handler — the next except up the stack — treats that as fatal, marks the engine dead, and re-raises:
python vllm/v1/engine/core.py Lines 1229-1235 except Exception as e: if enginecore is None: logger.exception("EngineCore failed to start.") else: logger.exception("EngineCore encountered a fatal error.") enginecore.sendenginedead() raise e
Impact
Availability only. cachesalt is an attacker-controlled, publicly reachable request field that vLLM validates too loosely. A value such as "/" passes vLLM's check, reaches the stricter downstream validator, and its ValueError is never converted into a request-scoped failure — instead it kills the EngineCore process, a denial of service for every concurrent request on that server (HTTP failures, then /health failing).
Applicability: the built-in LMCache-MP KV connector must be enabled (lmcache >= 0.4.4), which is itself an opt-in KV-connector boundary. On such deployments no other special configuration is required, and the crash is a resource-availability failure rather than expected behavior of the opt-in feature.
Suggested Fix
Two independent fixes:
1. Tighten admission — add one shared validatecachesalt() helper (for example in vllm/entrypoints/openai/engine/protocol.py) that matches or exceeds the downstream IPCCacheServerKey rules — reject @, /, \, NUL, and >128-character salts at the HTTP boundary with a 4xx — and route every request model that exposes cachesalt through it: the three checkcachesaltsupport validators above, the pooling base request, and the token-in-token-out GenerateRequest (which has no validator today). 2. Defense in depth — wrap the LMCache-MP lookup call reached from Scheduler.schedule() so a downstream validator ValueError becomes a request-scoped failure instead of an EngineCore-fatal exception. This is the fix that also covers future divergence between vLLM's and LMCache's salt rules; the right failure semantics (fail the request vs. fall back to a cold lookup) is a maintainer design call.
The core of fix 1 is a single shared helper; each checkcachesaltsupport body then becomes validatecachesalt(data.get("cachesalt")), and GenerateRequest gains an equivalent mode="before" validator:
python vllm/entrypoints/openai/engine/protocol.py — new shared helper CACHESALTFORBIDDENCHARS = frozenset("@/\\\x00") MAXCACHESALTLENGTH = 128
def validatecachesalt(cachesalt: object) -> None: """Validate cache salts before they reach downstream cache backends.""" if cachesalt is None: return if not isinstance(cachesalt, str) or not cachesalt: raise VLLMValidationError( "Parameter 'cachesalt' must be a non-empty string if provided.", parameter="cachesalt", ) if len(cachesalt) > MAXCACHESALTLENGTH or any( char in CACHESALTFORBIDDENCHARS for char in cachesalt ): raise VLLMValidationError( "Parameter 'cachesalt' must be at most 128 characters and must " "not contain '@', '/', '\\\\', or NUL.", parameter="cachesalt", )
This distinguishes the finding from GHSA-6qc9-v4r8-22xg: that advisory fixed only the guidedjson/xgrammar trigger of the same schedule()-into-runenginecore fatal-handler family, so the cachesalt admission gap and the schedule()-level defense-in-depth (fix 2) survive its published fix. A patch implementing fix 1 across all five request models, with a regression test covering the rejected ("/", 129-char) and accepted salt shapes, applies to v0.25.1 with line offsets and no fuzz.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
---
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51444
Other sources
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, OpenAI-compatible request models accept a non-empty cachesalt value without enforcing the character and length restrictions required by the IPCCacheServerKey consumer in LMCache-MP. On deployments using the LMCache-MP connector, a salt that contains a forbidden character or exceeds the permitted length can raise an uncaught ValueError during scheduler cache lookup, causing EngineCore to terminate and denying service to all concurrent users. This issue is fixed in version 0.30.0.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.30.0
Event History
Frequently Asked Questions
Which deployments are exposed to this denial of service?
The issue affects vLLM versions before 0.30.0 when they use the LMCache-MP connector. The failure occurs in the IPCCacheServerKey cache-key consumer during scheduler cache lookup.
What access does an attacker need to trigger the failure?
An attacker needs privileges to submit an OpenAI-compatible request to the affected vLLM service. A single request with a non-empty cache_salt containing a forbidden character or exceeding the permitted length can terminate EngineCore.
Is service impact limited to the requesting user?
No. EngineCore termination denies service to all concurrent users of the affected deployment.
What version resolves the issue?
Upgrade vLLM to version 0.30.0, which fixes the missing cache_salt validation.