GHSA-ph72-cqr5-qpp7: Input Validation

Published Oct 5, 2026
·
Updated

Affected

- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches.

Summary

vLLM's disaggregated scale-out transport splits a multimodal request into a trusted render step (POST /v1/chat/completions/render) and a separate generate step (POST /inference/v1/generate). The generate route decodes a caller-supplied features object — serialized encoder tensors (kwargsdata), multimodal hashes (mmhashes), placeholder ranges (mmplaceholders), and the internal field-processor selection — and forwards it into the engine as if it had come from the trusted renderer, with no rebinding to (or validation against) the active model's renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the /inference prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field.

Depending on which field is forged, this produces:

- an engine-fatal crash of the shared EngineCore process (denial of service), reproduced as a CUDA illegal-memory-access, a post-admission rank-mismatch ValueError, and a hard assert — three independent forged fields (sites 1, 2, 3); - silent cross-request encoder-cache poisoning / disclosure of shared encoder state when the cache-key hash is not bound to the payload (site 4); - transport-level integrity loss when sparse placeholder masks are dropped during render-to-generate replay (site 5).

All five share one root cause and one fix shape: the reconstructed multimodal state on the scale-out path is trusted without being rebound to, and validated against, the active model's renderer output before it reaches the engine.

These sites are distinct from prior multimodal hardening. Site 1 survives GHSA-wv77-2vpf-vmmg (that fix validates full tensor shape in MultiModalDataParser/getinputembeddings on the prompt-embeds path), because our request forges imagegridthw metadata with the pixel bytes intact and reaches the Qwen2 vision RoPE/cuseqlens and imageembeds.split sink, which the shape-check fix does not rebind. Site 4 is distinct from GHSA-c65p-x677-fgj6 (which folds metadata into MultiModalHasher.serializeitem to stop hash collisions), because the scale-out generate path trusts a caller-supplied mmhash as the cache key with no origin binding, so that fix does not stop a caller from submitting a victim's hash or a kwargsdata=None cache read.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1).

Shared entry point and control surface for all five sites:

- POST /inference/v1/generate route: vllm/entrypoints/scaleout/tokenintokenout/apirouter.py#L46-L75. - Public features schema (kwargsdata, mmhashes, mmplaceholders): vllm/entrypoints/scaleout/tokenintokenout/protocol.py#L42-L63. - ServingTokens.servetokens() copies the decoded geometry and hashes into engine structures without rebinding to the renderer schema: vllm/entrypoints/scaleout/tokenintokenout/serving.py#L145-L172. - The scale-out routers are registered by default on generate-capable servers and /inference is treated as an ordinary auth-guarded prefix: vllm/entrypoints/openai/apiserver.py#L217-L219.

Site 1 — forged Qwen grid geometry (engine-fatal DoS). The decoded imagegridthw is never rebound to the rendered pixel-tensor element count.

- Reconstruction with no model-specific geometry invariant: serving.py#L147-L172. - Qwen2VisionTransformer.prepareencodermetadata() derives RoPE tables, cuseqlens, and the FlashAttention maxseqlen from the caller-supplied grid: qwen2vl.py#L658-L712, consumed in forward() (#L713-L751). - processimageinput() computes imageembeds.split(sizes) from the same untrusted metadata: qwen2vl.py#L1348-L1369; M-RoPE positions at qwen2vl.py#L1223.

The only shape check is assert gridthw.ndim == 2; the split sizes and the vision-encoder call are then derived directly from the caller-supplied grid, with no cross-check against the pixel-tensor row count:

python vllm/modelexecutor/models/qwen2vl.py Lines 1348-1369 def processimageinput( self, imageinput: Qwen2VLImageInputs ) -> tuple[torch.Tensor, ...]: gridthw = imageinput["imagegridthw"] assert gridthw.ndim == 2

if imageinput["type"] == "imageembeds": imageembeds = imageinput["imageembeds"] else: pixelvalues = imageinput["pixelvalues"]

if self.usedataparallel: return rundpshardedmropevisionmodel( self.visual, pixelvalues, gridthw.tolist(), ropetype="rope3d" ) else: imageembeds = self.visual(pixelvalues, gridthw=gridthw)

# Split concatenated embeddings for each image item. mergesize = self.visual.spatialmergesize sizes = (gridthw.prod(-1) // mergesize // mergesize).tolist() return imageembeds.split(sizes)

Site 2 — wire-selected field-processor type confusion (engine-fatal DoS). MsgpackDecoder.decodemmfieldelem() trusts a wire-selected field-factory name and constructs the internal field processor directly from caller data.

- vllm/v1/serialutils.py#L440-L454 reads factorymethname, factorykw = obj["field"] and calls getattr(MultiModalFieldConfig, factorymethname). - Qwen2-VL field/schema contract: qwen2vl.py#L763-L791 and Qwen2VLImagePixelInputs (#L119-L144); parse/validate at qwen2vl.py#L1300. - Rank mismatch raised post-admission: vllm/utils/tensorschema.py#L155-L171, reached from TensorSchema.init → validate() (#L63).

Site 3 — non-positive placeholder length (engine-fatal DoS via reachable assert). PlaceholderRangeInfo{offset,length} is accepted as unconstrained integers and copied verbatim into the engine's PlaceholderRange.

- Schema: protocol.py#L28-L35. - Copied into PlaceholderRange: serving.py#L147-L153. - The only guard is an upper bound on embed count (numembeds > mmencodercachesize), with no non-positive check: vllm/v1/engine/inputprocessor.py#L459-L464; raw length returned by vllm/multimodal/inputs.py#L152-L154; window selection assumes non-empty ranges at vllm/multimodal/utils.py#L114-L134. - Sink: assert startidx < endidx at vllm/v1/worker/gpu/mm/encoderrunner.py#L114 (duplicated at vllm/v1/worker/gpumodelrunner.py#L3192); turned into a fatal shutdown by EngineCore's uncaught-exception path at vllm/v1/engine/core.py#L1229-L1233.

A length of 0 makes numencodertokens == 0, so endidx collapses to 0 and the bare assert fires inside the engine worker:

python vllm/v1/worker/gpu/mm/encoderrunner.py Lines 108-114 posinfo = mmfeature.mmposition startpos = posinfo.offset numencodertokens = posinfo.length

startidx = max(curquerystart - startpos, 0) endidx = min(curqueryend - startpos, numencodertokens) assert startidx < endidx

Site 4 — cache hash not bound to payload (integrity / disclosure). features.mmhashes (the cache key) and kwargsdata (the tensor) are independent fields with no origin or integrity binding.

- Schema exposing the caller-controlled hash and the None = resolve-from-cache semantics: protocol.py#L42-L63. - Handler forwards the caller's hashes unchanged into mminput(...): serving.py#L164-L170. - Input processing copies the caller hash into the feature identifier with only a string-type check: vllm/v1/engine/inputprocessor.py#L165-L181. - Sink — the receiver cache returns the cached tensor solely by that key (cachekey = feature.mmhash or feature.identifier): vllm/multimodal/cache.py#L602-L607.

The cache key is the caller-supplied hash with no verification against the tensor bytes, so a forged mmhash both stores under and reads back another request's slot:

python vllm/multimodal/cache.py Lines 601-607 for feature in mmfeatures: cachekey = feature.mmhash or feature.identifier self.touchreceivercacheitem(cachekey, feature.data)

for feature in mmfeatures: cachekey = feature.mmhash or feature.identifier feature.data = self.getandupdateitem(feature.data, cachekey) return mmfeatures

Site 5 — dropped sparse placeholder mask (transport integrity loss). The render path serializes placeholders as only offset/length, so models relying on sparse isembed masks lose the mask during render-to-generate replay.

- ServingRender.extractmmfeatures() builds each PlaceholderRangeInfo(offset=p.offset, length=p.length), discarding isembed: vllm/entrypoints/scaleout/render/serving.py#L212-L229. - Transport schema carries no field for the mask: protocol.py#L28. - ServingTokens.servetokens() reconstructs a dense PlaceholderRange regardless of the original: serving.py#L148-L152.

Impact

A single authenticated request to a scale-out multimodal deployment can:

- Crash the shared EngineCore process (sites 1, 2, 3), taking the served model down for every tenant (/health → 503). Availability-only; no code execution or data disclosure demonstrated for these sites. - Silently poison or read back another request's shared encoder-cache state (site 4) — an integrity/disclosure primitive. Attack complexity is High because the attacker must know or induce the victim's content hash; there is no availability impact for this site. - Corrupt backend-visible placeholder semantics across the render-to-generate boundary (site 5) for models that depend on sparse isembed masks.

The forged multimodal payload is small; only the trust in its self-declared geometry/identity is the defect.

Suggested Fix

On the scale-out path, do not trust caller-supplied multimodal state as renderer-produced. After decoding features, rebind and validate the reconstructed MultiModalKwargsItem against the active model's renderer contract at the HTTP boundary:

1. Reject any request whose decoded grid geometry is inconsistent with the rendered pixel-tensor element count and declared placeholder span, before it reaches prepareencodermetadata() (site 1). 2. Rebind each field's processor type to the schema the active model's renderer declares (or reject if it differs), instead of reconstructing internal field processors from wire-selected factory names (site 2). 3. Reject any PlaceholderRangeInfo with length <= 0 (or out-of-range offset) with a request-scoped 4xx, and convert the encoder-runner invariant into a checked, request-scoped error rather than a process-fatal assert (site 3). 4. Recompute or verify the content hash for submitted kwargsdata before using it as a cache key, and namespace receiver-cache keys to a server-generated or principal scope, refusing cache-reads for hashes the caller did not legitimately produce (site 4). 5. Serialize isembed in PlaceholderRangeInfo, validate its length against the placeholder span, and reconstruct it on replay (site 5).

Site 1 — validate grid geometry before the vision encoder. Replace the bare assert gridthw.ndim == 2 in processimageinput()/processvideoinput() with a shared helper that recomputes the split sizes and rejects a grid whose patch-row count does not match the pixel tensor (and rejects non-positive / non-merge-divisible dims), so the mismatch never reaches imageembeds.split():

python vllm/modelexecutor/models/qwen2vl.py — processimageinput() - gridthw = imageinput["imagegridthw"] - assert gridthw.ndim == 2 + gridthw = imageinput["imagegridthw"] + inputtype = imageinput["type"] + inputtensor = ( + imageinput["imageembeds"] + if inputtype == "imageembeds" + else imageinput["pixelvalues"] + ) + sizes = validateqwen2vlinputgeometry( + modality="image", + inputtype=inputtype, + inputtensor=inputtensor, + gridthw=gridthw, + spatialmergesize=self.visual.spatialmergesize, + ) ... - # Split concatenated embeddings for each image item. - mergesize = self.visual.spatialmergesize - sizes = (gridthw.prod(-1) // mergesize // mergesize).tolist() return imageembeds.split(sizes)

where the helper raises before the encoder runs:

python vllm/modelexecutor/models/qwen2vl.py — new validateqwen2vlinputgeometry() if t <= 0 or h <= 0 or w <= 0: raise ValueError(f"{modality} gridthw row {index} must be positive ...") if h % spatialmergesize != 0 or w % spatialmergesize != 0: raise ValueError(f"{modality} gridthw row {index} must be divisible ...") ... if actualrows != expectedrows: raise ValueError( f"{modality} {rowkind} do not match gridthw: " f"expected {expectedrows}, got {actualrows}." )

Site 3 — constrain the placeholder schema. Make PlaceholderRangeInfo reject non-positive lengths and negative offsets at the Pydantic boundary (plus parallel-length, non-overlapping, and within-prompt validators), turning the process-fatal assert into a request-scoped 422:

python vllm/entrypoints/scaleout/tokenintokenout/protocol.py — PlaceholderRangeInfo - offset: int - length: int + offset: int = Field(ge=0) + length: int = Field(gt=0)

Site 4 — bind the cache key to the payload. Derive each cache key by hashing the submitted serialized tensor (ignoring the caller's mmhashes) and refuse cache-only reads, so a forged hash can neither poison nor read a victim's slot:

python vllm/entrypoints/scaleout/tokenintokenout/serving.py — servetokens() + mmhashes = bindmmhashestokwargsdata( + features.mmhashes, features.kwargsdata, + ) engineinput = mminput( prompttokenids=request.tokenids, mmkwargs=MultiModalKwargsItems(mmkwargs), - mmhashes=features.mmhashes, + mmhashes=mmhashes, mmplaceholders=mmplaceholders, cachesalt=request.cachesalt, )

where bindmmhashestokwargsdata() raises on kwargsdata is None (cache-only read) and derives sha256(modality || "\0" || serializeditem) per item. Site 2 applies the same rebind-to-declared-schema pattern in mmserde.py (passing modality + mmprocessor into decodemmkwargsitem), and site 5 adds an isembed field to PlaceholderRangeInfo with fromplaceholderrange/toplaceholderrange helpers so the sparse mask survives render-to-generate replay. Each site fix ships with a regression test. This packet groups the five sites because they share one entry point (/inference/v1/generate + /render) and one root cause; we are happy to split it into per-component advisories (for example, engine-fatal input-validation vs. cache-key binding vs. transport-schema integrity) if the vLLM team prefers.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

These vulnerabilities were discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.

---

Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51898

Affected Software

1 affected componentFixes available
pip/vllm<0.30.0
0.30.0

Remediation

Recommended actions to resolve this vulnerability, in priority order.

  1. Upgrade

    Upgrade pip/vllm to a version that resolves this vulnerability.

    Fixed in 0.30.0
  2. Compensating control

    In vLLM Qwen2-VL input processing, validate decoded image_grid_thw before the vision encoder: require positive dimensions, require height and width to be divisible by spatial_merge_size, recompute the split sizes, and reject any grid whose patch-row count does not match the rendered pixel-tensor element count or declared placeholder span.

  3. Compensating control

    On the scale-out multimodal decode path, rebind each field's processor type to the schema declared by the active model's renderer, or reject the request if it differs; do not construct internal field processors from caller-selected factory_meth_name values.

  4. Compensating control

    Constrain PlaceholderRangeInfo at the request boundary by requiring length > 0 and offset >= 0, and validate parallel lengths, non-overlap, and containment within the prompt; convert the encoder-runner start_idx < end_idx invariant into a request-scoped validation error instead of a process-fatal assert.

  5. Compensating control

    Bind multimodal cache keys to the submitted kwargs_data by deriving each key from the serialized tensor, such as sha256(modality || "\0" || serialized_item); ignore caller-supplied mm_hashes for key derivation, refuse kwargs_data=None cache-only reads, and namespace receiver-cache keys to a server-generated or principal scope.

  6. Compensating control

    Add is_embed to PlaceholderRangeInfo, validate its length against the placeholder span, and preserve and reconstruct the sparse mask across render-to-generate replay instead of serializing only offset and length.

Event History

Oct 5, 2026
Advisory Published
via GitHub·11:42 PM
Data Sourced
via GitHub·11:42 PM
DescriptionSeverityWeaknessAffected Software

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203