GHSA-58v5-2m8f-94pr: Medium severity pip/vllm vulnerability
Summary
The OpenAI-compatible chat endpoint accepts request-level video loader options through mediaiokwargs. A caller can select the GLMGA sampler and provide large fps and maxframes values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests.
The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed.
Affected configuration
The server must expose chat completions for a video-capable model and accept request-level mediaiokwargs. The model does not need to use GLMGA by default because the request value overrides the model's loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required.
The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, any caller holding a valid key can reach the same path.
Attack surface
A remote caller submits a valid chat-completion request with a small video and the following request-level options:
json { "mediaiokwargs": { "video": { "videobackend": "glmga", "backend": "opencv", "fps": 500000, "maxframes": 500000 } } }
Increasing both numeric values increases temporary list construction and deduplication work even when the final decoded frame set remains unchanged.
Root cause
1. ChatCompletionRequest exposes mediaiokwargs as request-controlled nested values: protocol.py#L365-L371. 2. The request options are carried into chat parameters without a numeric work bound on GLMGA's fps or maxframes: protocol.py#L571-L599. 3. Media options are merged so request values override configured defaults, including videobackend: connector.py#L577-L605. 4. The connector selects a registered video loader and runs media loading through the shared executor: video.py#L28-L66 and connector.py#L44-L47. 5. GLMGA calculates extractt = min(int(duration fps), maxframes), builds a list with that many entries, and deduplicates it before decoding frames: video.py#L667-L740. 6. The OpenCV decoder receives the already-deduplicated indices, so a two-frame file can trigger a large intermediate allocation while producing only two decoded frames: opencv.py#L28-L50.
The missing invariant is a strict upper bound on sampling work before the index list is constructed. Bounding decoded output after deduplication does not bound the vulnerable intermediate computation.
Suggested remediation
Validate request-level video sampling options before dispatching work to the media executor. Enforce conservative absolute limits for fps, maxframes, and especially the computed candidate count. Reject non-finite, negative, or otherwise invalid numeric values.
Avoid constructing O(extractt) intermediate Python lists. Generate bounded unique frame indices directly from the source frame count and output-frame limit. Apply the limit before list construction and before shared executor submission. Add tests showing that extreme request values are rejected or consume constant memory for a fixed output-frame count.
Workarounds
- Remove or filter request-level videobackend, fps, and maxframes options at the gateway. - Do not allow untrusted callers to select GLMGA. - Apply authentication, rate limiting, request concurrency limits, and process memory isolation. - Use a separate constrained media-loading worker pool where operationally feasible. - Existing media byte, pixel, or decoded-frame limits do not necessarily bound this pre-decode candidate-index allocation.
Related advisory
GHSA-cqm8-jxg6-fqfq concerns partial denial of service in the DeepStream video backend through backend confusion and insufficient pixel guarding. This report concerns a different CPU-side GLMGA algorithm: request-controlled sampling values create a large redundant index list before the OpenCV decoder reads a fixed, tiny frame set. The vulnerable backend, resource sink, and required remediation differ.
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.30.0 - Compensating control
Require authentication, apply rate limiting and request-concurrency limits, and isolate media-processing memory for video requests.
- Compensating control
At the gateway, remove or filter request-level video_backend, fps, and max_frames options, and do not allow untrusted callers to select the GLMGA backend.
- Compensating control
Validate request-level video sampling options before dispatching work to the media executor; reject non-finite, negative, or otherwise invalid fps and max_frames values, and enforce conservative absolute limits on fps, max_frames, and the computed candidate count before list construction.
- Compensating control
Generate bounded unique frame indices directly from the source frame count and output-frame limit instead of constructing O(extract_t) intermediate Python lists and deduplicating them.
- Compensating control
Use a separate constrained media-loading worker pool for video processing where operationally feasible.