GHSA-x6mc-67gf-chw4: Medium severity pip/vllm vulnerability
Summary
An unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level mediaiokwargs.video.maxframes and fps knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated POST /tokenize.
The numframes ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read numframes at all. The same knobs were already capped upstream for GLMGAVideoBackend as an accepted security fix in 8b6de0eb9 (PR #54935, merged 2026-09-04); that cap never reached Qwen.
Details
Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first)
GHSA-vxqj-p4gw-9h4c reported that request-level mediaiokwargs.video.numframes overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping numframes inside VideoMediaIO.mergekwargs.
That clamp does not reach the Qwen samplers. Qwen2VLVideoBackend and Qwen3VLVideoBackend do not read numframes at all — Qwen2VLVideoBackend's own docstring says so ("numframes is ignored (fps-driven, like the Qwen3-VL loader)"). They bound on maxframes, read from the same merged dict with no ceiling:
python vllm/multimodal/video.py — Qwen3VLVideoBackend.computeframesindextosample minframes = kwargs.get("minframes", 4) maxframes = kwargs.get("maxframes", 768) numframes = int(totalframesnum / originalfps fps) numframes = min(max(numframes, minframes), maxframes, totalframesnum)
With maxframes raised from the request, the only remaining bound is totalframesnum — every frame in the container.
I applied PR #51969's patch locally and re-ran both paths through the real merge layer (mergemediaiokwargs → VideoMediaIO.mergekwargs → MediaConnector), against main @ b23433088:
| request mediaiokwargs.video | merged kwargs after #51969 | frames decoded | peak RSS | |---|---|---|---| | (absent) | None | 32 | 706 MiB | | {"videobackend":"opencv","numframes":-1} | {…,"numframes":32} | 32 — fixed | 706 MiB | | {"videobackend":"qwen3vl"} | {…,"numframes":32} | 60 | 779 MiB | | {"videobackend":"qwen3vl","maxframes":1e9,"fps":1e6} | {…,"maxframes":1000000000,"fps":1000000,"numframes":32} | 900 — survives | 2 994 MiB |
#51969 does exactly what it claims for numframes; the clamp writes numframes: 32 into the merged dict and the Qwen sampler ignores it, while maxframes and fps pass through untouched.
The codebase already has the fix pattern, on other backends
This is not a new control being proposed. Commit 8b6de0eb9 — "[Security] Cap GLMGA video sampling to prevent request-driven resource exhaustion" (PR #54935, merged 2026-09-04, same author as #51969) — caps precisely these two knobs for GLMGAVideoBackend:
python MAXFRAMES: ClassVar[int] = 640 MAXFPS: ClassVar[int] = 30 ... targetfps = min(target.fps, cls.MAXFPS) maxframes = min(kwargs.get("maxframes", cls.MAXFRAMES), cls.MAXFRAMES)
Its description states the root cause as "both targetfps and maxframes are controllable via request-level mediaiokwargs", and says it follows "the pattern established by GLM46VVideoBackend" (which caps via MAXFRAMECOUNTDYNAMIC = 640 and MAXDURATION = 2400).
So two backends cap request-controlled fps/maxframes as an accepted security measure. Qwen2VLVideoBackend and Qwen3VLVideoBackend — the most widely deployed video models on vLLM — cap neither. In GLMGA the uncapped knobs sized an intermediate index list; in the Qwen samplers they size the decoded frame buffer, which is larger by the per-frame pixel count.
Reachability — default configuration, no authentication
mediaiokwargs is a request body field on ChatCompletionRequest and is carried by /v1/chat/completions, /v1/embeddings, /v1/responses, /tokenize and /invocations. No flag gates it. apikey defaults to None (vllm/entrypoints/launchers/cliargs.py), and AuthenticationMiddleware is installed only when a key is configured — a default vllm serve is entirely unauthenticated. Even with --api-key set, GUARDEDPREFIX = ("/v1", "/v2", "/inference", "/cohere") (vllm/entrypoints/serve/middleware/authenticate.py:11), so /invocations — which validates the same ChatCompletionRequest body — and /tokenize — which performs full media ingestion — remain unauthenticated. MediaConnector.fetchvideo applies the model's registered sampler only when the request did not name one (if "videobackend" not in videoiokwargs), so videobackend: "qwen3vl" is selectable on any deployment; request-level selection of stock sampler subclasses is already established as reachable by GHSA-j682-9xp5-rrf3. On a Qwen deployment no videobackend key is needed at all. Default-on for any video-capable Qwen model.
The Rust frontend is not affected. rust/src/server/src/routes/openai/chatcompletions/validate.rs:90 rejects mediaiokwargs with "mediaiokwargs is not supported." No second front is needed.
Impact
Unauthenticated remote denial of service by memory exhaustion of the API-server process. The decode runs in the frontend during chat parsing, before scheduling or admission control, so every tenant on the instance is affected. --limit-mm-per-prompt does not apply — it bounds media items, not frames within an item. CWE-770 / CWE-400.
Measured over unauthenticated HTTP: 1.43 MiB request body, 74 extra bytes of JSON, server peak RSS 2 271 → 13 629 MiB. In-process, the same request takes the decode from 60 to 900 frames. The frame count is bounded only by totalframesnum, which is the attacker's choice of video, and decoded bytes are frames × H × W × 3.
Context, measured on the numframes path (GHSA-vxqj-p4gw-9h4c's path), not this one — these figures show what an unbounded frame count costs once the source video is chosen for it, and they transfer to this path because both converge on the same readframesnorecovery allocation:
5.74 MiB request body → 9.27 GiB decoded, 9.73 GiB of new resident memory (1 736×). 117.6 MiB video/jpeg payload → process OOM-killed: Out of memory: Killed process 182667 (python) total-vm:38665704kB, anon-rss:21121616kB
The two were not re-run at the larger sizes through the Qwen sampler; the 900-frame figure above is what I measured on this path.
Suggested fix
Extend the ceiling to the sampler-side knobs rather than clamping the single numframes key. Two options, either acceptable:
1. Per-backend caps, matching 8b6de0eb9. Give Qwen2VLVideoBackend and Qwen3VLVideoBackend the MAXFRAMES / MAXFPS treatment already applied to GLMGAVideoBackend, so kwargs.get("maxframes", …) and target.fps are clamped to class constants. Smallest change; consistent with the accepted precedent. It leaves Glm5NextVideoBackend, Molmo2VideoBackend, NemotronVLVideoBackend, DynamicVideoBackend and OpenCVDynamicOpenPanguVideoBackend to be audited one by one, which is the current trajectory (#55727, #56207, #56390). 2. Strip the frame-count knobs at the merge boundary. In VideoMediaIO.mergekwargs, drop maxframes / minframes / fps from runtimekwargs the way hwdecoders, poolsize, device and unconfigured GPU backends are already dropped there. That treats the whole frame-count family as startup-only in one place and is robust to future sampler subclasses, at the cost of removing a request-level knob some users may rely on. A clamp-don't-strip variant (request may lower, never raise) preserves the feature.
Option 2 composes with #51969 and needs no per-backend audit; I would favour it, but option 1 is the more conservative change and matches what has already been merged.
Affected versions
>= 0.24.0, through v0.29.1rc0 and main @ b23433088.
Lower bound established by probing release tags through the GitHub contents API; no version below is inferred, each was read out of the file at that tag.
Confirmed present — Qwen samplers reading unclamped maxframes (maxframes = kwargs.get("maxframes", 768) inside Qwen2VLVideoBackend / Qwen3VLVideoBackend in vllm/multimodal/video.py):
| ref | Qwen3VLVideoBackend present | unclamped maxframes | |---|---|---| | v0.23.0 | no (class does not exist) | n/a | | v0.24.0 | yes | yes | | v0.25.0, v0.26.0, v0.27.0, v0.28.0, v0.29.0, v0.29.1rc0 | yes | yes | | main @ b23433088 | yes | yes |
Confirmed present — request-level mediaiokwargs (field on the chat request model, and VideoMediaIO.mergekwargs present): every ref probed, v0.19.0 through v0.29.1rc0 and main @ b23433088.
Confirmed absent — any numframes ceiling clamp (PR #51969 unmerged): every ref probed, v0.19.0 through v0.29.1rc0 and main @ b23433088.
Not resolved, and why:
1. The introducing commit/PR for Qwen2VLVideoBackend / Qwen3VLVideoBackend. My clone is shallow (--depth=300), so git log -S'maxframes' cannot reach it; the v0.23.0 → v0.24.0 boundary above is the tightest bound I established by probing release tags. 2. The rc tags between v0.23.0 and v0.24.0, to tighten the lower bound to a specific release candidate. 3. Whether a maxframes knob on a differently named pre-v0.24.0 backend is separately affected — at v0.23.0 GLMGAVideoBackend already carried maxframes = kwargs.get("maxframes", 640) with no clamp, and that clamp was only added on 2026-09-04 by 8b6de0eb9. Versions between are plausibly affected through GLMGA rather than Qwen; I did not test that path. 4. Whether GHSA-vxqj-p4gw-9h4c's own affected range differs, which I cannot see — the advisory returns 404 to me.
No dependency versions are asserted anywhere in this report.
Weaknesses of this report, stated upfront
No GPU was used. This host has no CUDA device, so the HTTP results come from the in-tree GPU-less render server rather than a full vllm serve. That server runs the real frontend — the same request model, the same mergemediaiokwargs → VideoMediaIO.mergekwargs → MediaConnector.fetchvideo chain, the same /tokenize route — and media decoding happens entirely in the frontend, so I do not believe the engine's presence changes the result. I have not confirmed that on a GPU deployment, and a reviewer may reasonably want that repeated under vllm serve. The in-process measurements (the #51969 comparison table) were taken by importing the tree directly at b23433088, verified by file, with no vllm wheel installed. Amplification figures are a floor. The OpenCV build available here offers only mp4v/XVID, giving ~2 200× compression on static content. An attacker using x264/x265 would do materially better for the same frame count. The two large figures in the Impact section (9.73 GiB RSS; the OOM kill) were measured on the numframes path, not this one, and are labelled as such. Per-frame pixels remain bounded by VLLMMAXIMAGEPIXELS; this concerns the unbounded frame count, which multiplies it. VLLMMAXMEDIADOWNLOADSIZEMB (default 256) caps the compressed size of an HTTP-fetched video but not a data: URI, which MediaConnector.loadfromurl dispatches before any size logic is reached.
---
Prepared with AI assistance (Claude), per the repository's contributing guidance on disclosing AI-assisted contributions. All findings were verified by executing vLLM's own code at the commits cited; the analysis and the claims are my own.
Reported by Eva Crystal / 0xiviel (XSource Security).
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.30.0 - Compensating control
In VideoMediaIO.merge_kwargs, strip max_frames, min_frames, and fps from runtime_kwargs at the merge boundary so request-level media_io_kwargs cannot control Qwen video frame sampling; treat these knobs as startup-only.