CVE-2026-105753: vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core
Affected
- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.
Summary
vLLM's default multimodal cache (mmprocessorcachetype="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that getandupdate() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1.
That invariant breaks when a request is rejected after P0 has rendered and hashed the multimodal input (populating the P0 cache) but before P1 receives the item — for example, an oversized chat prompt rejected on maxmodellen after rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends None instead of the payload, and P1 — which has nothing cached — trips assert mmitem is not None, f"Expected a cached item for {mmhash=}".
This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
- P0 metadata cache — MultiModalProcessorSenderCache at vllm/multimodal/cache.py#L379; getandupdateitem at #L410-L421, commit assertion at #L418. - P1 payload cache (the actual sink) — MultiModalReceiverCache at vllm/multimodal/cache.py#L630; getandupdateitem at #L652-L663, with the failing assert mmitem is not None, f"Expected a cached item for {mmhash=}" at #L660. - Default mmprocessorcachetype = "lru" at vllm/config/multimodal.py#L132; dispatch in vllm/multimodal/registry.py#L294-L307 (sender) and #L322-L331 (receiver). The processoronly, disabled-caching, and shm paths are not affected. - Rejection-after-render window — rendering happens before length validation in vllm/entrypoints/openai/chatcompletion/serving.py#L206-L231 (renderchatrequest), and the length check raises after the render in vllm/entrypoints/serve/utils/apiutils.py#L171-L189. - Cache-commit call site — vllm/multimodal/processing/processor.py#L1347 (mergemmkwargs) commits the P0 sender entry during render. - The separate stale-order eviction hang is already fixed via vllm/utils/cache.py#L120-L121 (LRUCache.touch() guarding if key in self:) — a different mechanism that does not touch the sender/receiver commit-ordering protocol and does not remediate this assertion.
The P1 receiver sink — the assertion that fires when P1 receives None for a hash it never cached:
python vllm/multimodal/cache.py Lines 651-663 @override def getandupdateitem( self, mmitem: MultiModalKwargsItem | None, mmhash: str, ) -> MultiModalKwargsItem: if (cacheditem := self.cache.get(mmhash)) is not None: return cacheditem
assert mmitem is not None, f"Expected a cached item for {mmhash=}"
self.cache[mmhash] = mmitem return mmitem
The P0 sender — on a hit it drops the payload (returns None) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:
python vllm/multimodal/cache.py Lines 409-422 @override def getandupdateitem( self, mmitem: MultiModalProcessorCacheInItem, mmhash: str, ) -> MultiModalProcessorCacheOutItem: if (cacheditem := self.cache.get(mmhash)) is not None: return None, cacheditem.promptupdates
assert mmitem is not None, f"Expected a cached item for {mmhash=}"
self.cache[mmhash] = MultiModalProcessorCacheItemMetadata(mmitem)
return mmitem
Impact
A remote client submitting multimodal requests can poison a cache identity — render a media item successfully, then have that request rejected — so a later request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.
On this revision the failure is scoped as a request-level preprocessing error (the engine core catches around preprocessaddrequest); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored lru cache.
Suggested Fix
Two complementary changes:
1. Make the mirrored commit atomic with admission — insert into the P0 sender cache only after the request has passed all admission checks (length, limits) and P1 has acknowledged the item, or roll back the P0 insert on rejection. 2. Defense in depth — convert the P1 receiver assert mmitem is not None into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.
The core of the rollback half: wrap the post-render length check so a ValueError rejection discards the P0 entries the render just committed, before re-raising. Add a discardsendercacheitem() on the processor cache (no-op default, pop on the sender) and a Renderer.discardmmcacheentries() that walks a rendered request's mmhashes:
python vllm/entrypoints/openai/chatcompletion/serving.py (createchatcompletion) - maxtokens = getmaxtokens( - maxmodellen, - ..., - truncateprompttokens=request.truncateprompttokens, - ) + try: + maxtokens = getmaxtokens( + maxmodellen, + ..., + truncateprompttokens=request.truncateprompttokens, + ) + except ValueError: + for renderedinput in engineinputs: + if mmhashes := renderedinput.get("mmhashes"): + self.renderer.discardmmcacheentries(mmhashes) + raise
python vllm/multimodal/cache.py (MultiModalProcessorSenderCache) + @override + def discardsendercacheitem(self, mmhash: str) -> None: + self.cache.pop(mmhash, None)
This closes the maxmodellen rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 assert a checked, request-scoped error) is recommended.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
---
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897
Other sources
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receiver cache never receives the payload if that request is rejected. A later request reusing the same media hash causes MultiModalProcessorSenderCache to send no payload and MultiModalReceiverCache to reach an assertion with the message "Expected a cached item," producing a shared-service availability failure. This issue is fixed in version 0.28.0.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.28.0 - Upgrade
Upgrade
vllmto a version that resolves this vulnerability.Fixed in 0.28.0
Event History
Frequently Asked Questions
Who is exposed to this availability failure?
vLLM deployments prior to 0.28.0 using the default mirrored multimodal LRU cache are affected. The impact is a shared-service availability failure when the cache desynchronization is triggered.
What must happen for an attacker to trigger the assertion?
An attacker needs permission to submit requests. A multimodal request must be rejected after its media hash has been committed to the frontend sender cache, then a later request must reuse that same media hash; the later request can cause the receiver assertion.
Are default configurations affected?
Yes. The affected cache is described as the default mirrored multimodal LRU cache.
How can I recognize that this issue has occurred?
The engine receiver cache reaches an assertion with the message "Expected a cached item." This occurs when the sender cache omits a payload for a reused media hash that was never received by the engine cache.
What version fixes the issue?
Upgrade vLLM to version 0.28.0, which fixes this issue.