CVE-2026-105760: vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
Summary
The OpenAI-compatible chat endpoint accepts request-level video loader options through mediaiokwargs. A caller can select the GLMGA sampler and provide large fps and maxframes values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests.
The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed.
Affected configuration
The server must expose chat completions for a video-capable model and accept request-level mediaiokwargs. The model does not need to use GLMGA by default because the request value overrides the model's loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required.
The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, any caller holding a valid key can reach the same path.
Attack surface
A remote caller submits a valid chat-completion request with a small video and the following request-level options:
json { "mediaiokwargs": { "video": { "videobackend": "glmga", "backend": "opencv", "fps": 500000, "maxframes": 500000 } } }
Increasing both numeric values increases temporary list construction and deduplication work even when the final decoded frame set remains unchanged.
Root cause
1. ChatCompletionRequest exposes mediaiokwargs as request-controlled nested values: protocol.py#L365-L371. 2. The request options are carried into chat parameters without a numeric work bound on GLMGA's fps or maxframes: protocol.py#L571-L599. 3. Media options are merged so request values override configured defaults, including videobackend: connector.py#L577-L605. 4. The connector selects a registered video loader and runs media loading through the shared executor: video.py#L28-L66 and connector.py#L44-L47. 5. GLMGA calculates extractt = min(int(duration fps), maxframes), builds a list with that many entries, and deduplicates it before decoding frames: video.py#L667-L740. 6. The OpenCV decoder receives the already-deduplicated indices, so a two-frame file can trigger a large intermediate allocation while producing only two decoded frames: opencv.py#L28-L50.
The missing invariant is a strict upper bound on sampling work before the index list is constructed. Bounding decoded output after deduplication does not bound the vulnerable intermediate computation.
Suggested remediation
Validate request-level video sampling options before dispatching work to the media executor. Enforce conservative absolute limits for fps, maxframes, and especially the computed candidate count. Reject non-finite, negative, or otherwise invalid numeric values.
Avoid constructing O(extractt) intermediate Python lists. Generate bounded unique frame indices directly from the source frame count and output-frame limit. Apply the limit before list construction and before shared executor submission. Add tests showing that extreme request values are rejected or consume constant memory for a fixed output-frame count.
Workarounds
- Remove or filter request-level videobackend, fps, and maxframes options at the gateway. - Do not allow untrusted callers to select GLMGA. - Apply authentication, rate limiting, request concurrency limits, and process memory isolation. - Use a separate constrained media-loading worker pool where operationally feasible. - Existing media byte, pixel, or decoded-frame limits do not necessarily bound this pre-decode candidate-index allocation.
Related advisory
GHSA-cqm8-jxg6-fqfq concerns partial denial of service in the DeepStream video backend through backend confusion and insufficient pixel guarding. This report concerns a different CPU-side GLMGA algorithm: request-controlled sampling values create a large redundant index list before the OpenCV decoder reads a fixed, tiny frame set. The vulnerable backend, resource sink, and required remediation differ.
Other sources
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level mediaiokwargs field to select the GLMGA video backend and supply large values for the fps and maxframes options without a strict work ceiling. GLMGA constructs and deduplicates an attacker-sized pre-decode frame-index list, allowing a compact request and tiny valid video to consume disproportionate CPU time and memory in the shared media-loading executor. This issue is fixed in version 0.30.0.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.30.0 - Upgrade
Upgrade
vLLMto a version that resolves this vulnerability.Fixed in 0.30.0 - Compensating control
At the gateway, remove or filter request-level video_backend, fps, and max_frames options, and do not allow untrusted callers to select GLMGA.
- Compensating control
Validate request-level video sampling options before dispatching work to the media executor: reject non-finite, negative, or otherwise invalid values, enforce conservative absolute limits for fps, max_frames, and the computed candidate count, apply the limit before list construction and shared executor submission, and generate bounded unique frame indices directly instead of constructing O(extract_t) intermediate Python lists.
- Compensating control
Protect the media-loading path with authentication, rate limiting, request concurrency limits, and process memory isolation; where feasible, use a separate constrained media-loading worker pool.
Event History
Frequently Asked Questions
Who is exposed to this issue?
vLLM deployments running versions earlier than 0.30.0 are exposed if untrusted callers can submit requests that set request-level media_io_kwargs. The issue can affect the shared media-loading executor, so one malicious request may consume resources needed by other requests.
What does an attacker need to exploit it?
An attacker needs only network access to submit a request; no privileges or user interaction are required. They can select the GLMGA video backend and provide large fps and max_frames values, using even a tiny valid video.
Are default settings affected?
The vulnerable behavior requires a request that explicitly uses media_io_kwargs to select the GLMGA backend and set large fps and max_frames values. The provided information does not establish whether ordinary default request handling reaches this path.
What is the remediation?
Upgrade vLLM to version 0.30.0, which fixes the issue. If an upgrade cannot be performed immediately, restrict untrusted access to request-level media_io_kwargs and prevent requests from selecting GLMGA with attacker-controlled fps or max_frames values.