GHSA-ph3r-5jfg-f84f: Medium severity pip/vllm vulnerability

Published Oct 6, 2026
·
Updated

Affected

- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.

Summary

vLLM's default multimodal cache (mmprocessorcachetype="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that getandupdate() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1.

That invariant breaks when a request is rejected after P0 has rendered and hashed the multimodal input (populating the P0 cache) but before P1 receives the item — for example, an oversized chat prompt rejected on maxmodellen after rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends None instead of the payload, and P1 — which has nothing cached — trips assert mmitem is not None, f"Expected a cached item for {mmhash=}".

This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

- P0 metadata cache — MultiModalProcessorSenderCache at vllm/multimodal/cache.py#L379; getandupdateitem at #L410-L421, commit assertion at #L418. - P1 payload cache (the actual sink) — MultiModalReceiverCache at vllm/multimodal/cache.py#L630; getandupdateitem at #L652-L663, with the failing assert mmitem is not None, f"Expected a cached item for {mmhash=}" at #L660. - Default mmprocessorcachetype = "lru" at vllm/config/multimodal.py#L132; dispatch in vllm/multimodal/registry.py#L294-L307 (sender) and #L322-L331 (receiver). The processoronly, disabled-caching, and shm paths are not affected. - Rejection-after-render window — rendering happens before length validation in vllm/entrypoints/openai/chatcompletion/serving.py#L206-L231 (renderchatrequest), and the length check raises after the render in vllm/entrypoints/serve/utils/apiutils.py#L171-L189. - Cache-commit call site — vllm/multimodal/processing/processor.py#L1347 (mergemmkwargs) commits the P0 sender entry during render. - The separate stale-order eviction hang is already fixed via vllm/utils/cache.py#L120-L121 (LRUCache.touch() guarding if key in self:) — a different mechanism that does not touch the sender/receiver commit-ordering protocol and does not remediate this assertion.

The P1 receiver sink — the assertion that fires when P1 receives None for a hash it never cached:

python vllm/multimodal/cache.py Lines 651-663 @override def getandupdateitem( self, mmitem: MultiModalKwargsItem | None, mmhash: str, ) -> MultiModalKwargsItem: if (cacheditem := self.cache.get(mmhash)) is not None: return cacheditem

assert mmitem is not None, f"Expected a cached item for {mmhash=}"

self.cache[mmhash] = mmitem return mmitem

The P0 sender — on a hit it drops the payload (returns None) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:

python vllm/multimodal/cache.py Lines 409-422 @override def getandupdateitem( self, mmitem: MultiModalProcessorCacheInItem, mmhash: str, ) -> MultiModalProcessorCacheOutItem: if (cacheditem := self.cache.get(mmhash)) is not None: return None, cacheditem.promptupdates

assert mmitem is not None, f"Expected a cached item for {mmhash=}"

self.cache[mmhash] = MultiModalProcessorCacheItemMetadata(mmitem)

return mmitem

Impact

A remote client submitting multimodal requests can poison a cache identity — render a media item successfully, then have that request rejected — so a later request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.

On this revision the failure is scoped as a request-level preprocessing error (the engine core catches around preprocessaddrequest); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored lru cache.

Suggested Fix

Two complementary changes:

1. Make the mirrored commit atomic with admission — insert into the P0 sender cache only after the request has passed all admission checks (length, limits) and P1 has acknowledged the item, or roll back the P0 insert on rejection. 2. Defense in depth — convert the P1 receiver assert mmitem is not None into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.

The core of the rollback half: wrap the post-render length check so a ValueError rejection discards the P0 entries the render just committed, before re-raising. Add a discardsendercacheitem() on the processor cache (no-op default, pop on the sender) and a Renderer.discardmmcacheentries() that walks a rendered request's mmhashes:

python vllm/entrypoints/openai/chatcompletion/serving.py (createchatcompletion) - maxtokens = getmaxtokens( - maxmodellen, - ..., - truncateprompttokens=request.truncateprompttokens, - ) + try: + maxtokens = getmaxtokens( + maxmodellen, + ..., + truncateprompttokens=request.truncateprompttokens, + ) + except ValueError: + for renderedinput in engineinputs: + if mmhashes := renderedinput.get("mmhashes"): + self.renderer.discardmmcacheentries(mmhashes) + raise

python vllm/multimodal/cache.py (MultiModalProcessorSenderCache) + @override + def discardsendercacheitem(self, mmhash: str) -> None: + self.cache.pop(mmhash, None)

This closes the maxmodellen rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 assert a checked, request-scoped error) is recommended.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.

---

Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897

Affected Software

1 affected componentFixes available
pip/vllm<0.28.0
0.28.0

Remediation

Recommended actions to resolve this vulnerability, in priority order.

  1. Upgrade

    Upgrade pip/vllm to a version that resolves this vulnerability.

    Fixed in 0.28.0

Event History

Oct 6, 2026
Advisory Published
via GitHub·12:02 AM
Data Sourced
via GitHub·12:02 AM
DescriptionSeverityWeaknessAffected Software

Frequently Asked Questions

1

Which deployments are affected by this issue?

The affected package is pip vllm through version 0.25.1. Deployments using the default multimodal cache setting, mm_processor_cache_type="lru", use the mirrored frontend and engine-core cache design involved in the issue.

2

What does an attacker need to do to trigger the cache desynchronization?

An attacker needs low-privilege access and must cause a request to be rejected after the frontend has rendered and hashed multimodal input but before the engine core receives it. An oversized chat prompt rejected due to max_model_len is one stated example; a later request must reuse the same media hash.

3

How can operators identify likely exposure?

Check whether the deployment runs vLLM 0.25.1 or an earlier version and uses the default LRU multimodal processor cache. A relevant request pattern is a rejected multimodal request followed by another request that reuses the same media input or hash.

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203