CVE-2026-73557: vLLM: Incomplete CVE-2025-62164 remediation can be bypassed by concurrent prompt parts
Executive Summary
The follow-up protection for CVE-2025-62164 is incomplete at vLLM revision 26587f9519e22a5c4549ead7595ad9ca3229c4fd. It wraps serialized prompt-embedding reconstruction and dense conversion in torch.sparse.checksparsetensorinvariants(), but PyTorch 2.11.0 implements that context with save/enable/restore operations over process-global state. Two prompt-embedding parts in one /v1/chat/completions request are gathered concurrently on the event loop's default executor. When one context exits before the other loads its tensor, it can restore the global flag to False while the second part remains inside its guard.
In a deterministic run against hash-verified source from the affected revision, the actual target loader rejected an invalid sparse payload as a negative control. The frozen chat tracker then scheduled benign and malicious parts on distinct asyncio0 and asyncio1 threads. The benign context exited, the malicious loader observed the invariant flag disabled, and torch.load(weightsonly=True) reconstructed indices [[10], [10]] for a declared shape of [3, 3]. The run intercepted the target's todense() call before it operated on the invalid tensor.
This primary trigger requires --enable-prompt-embeds, which is default-off, but it does not require renderernumworkers > 1, a multimodal model, or --enable-mm-embeds. API authentication is optional in the stock server: middleware is installed only when CLI or environment API keys are supplied.
The lab proves bypass of the follow-up guard, invalid sparse reconstruction, and guarded-sink reachability. Crash and memory-corruption consequences are conditional on the behavior documented by the published CVE.
Background
CVE-2025-62164 / GHSA-mrw7-hf4f-83pf concerns client-controlled serialized promptembeds reaching torch.load(weightsonly=True) and an invalid sparse tensor reaching todense(). The advisory attributes memory corruption, denial of service, and potential code execution to that historical unsafe operation.
The remediation chronology matters for duplicate handling:
- PR #27204, merge commit 58fab50d82838d5014f4a14d991fdb9352c9c84b on 2025-10-22, introduced the default-off enablepromptembeds gate. It did not add the sparse-invariant context. - Commit 84e23d103d3483f944780d0d42bcf0993fd27e3a on 2025-12-15, titled additional protection for CVE-2025-62164 (#30649), added the process-global sparse-invariant context around load, type check, and dense conversion. - Refactor commit f0a1c8453ad1c664c8a04c83fe545195fcd556eb on 2026-01-31 moved the guarded loader into vllm/renderers/embedutils.py while preserving the same context. - Chat content-part commit 14043dfecd35dd2f12b4d51eb9fa166184a0ca0f on 2026-05-01 introduced promptembeds chat parts and the concurrent one-request schedule described here.
This report therefore does not present the malformed sparse payload or todense() sink as new. It reports a distinct concurrency root cause and trigger: unsynchronized save/enable/restore of the process-global follow-up guard, reachable through the later multi-part chat scheduler.
The affected revision pins PyTorch 2.11.0 in pyproject.toml:10.
Vulnerability Details
The target's safeloadpromptembeds performs the guarded operation in vllm/renderers/embedutils.py:16-39:
python with torch.sparse.checksparsetensorinvariants(): tensor = torch.load( BytesIO(pybase64.b64decode(embed, validate=True)), weightsonly=True, maplocation=torch.device("cpu"), ) if not isinstance(tensor, torch.Tensor): raise VLLMValidationError(...) tensor = tensor.todense()
The context is not request-local. With the global flag initially disabled, we can describe the verified interleaving:
1. Benign part A enters, saves False, and enables the flag. 2. Malicious part B enters, saves True, and leaves the flag enabled. 3. A completes its load and exits, restoring its saved False value. 4. B remains lexically inside its context but observes the actual global flag as False. 5. B's torch.load(..., weightsonly=True) reconstructs the malformed sparse tensor. 6. The target reaches tensor.todense() before later rank, hidden-size, and dtype checks.
weightsonly=True constrains deserialization types; it does not compensate for a sparse invariant check that another request has disabled.
The complete stock actor-to-sink chain, traced in the affected source, is:
POST /v1/chat/completions (vllm/entrypoints/openai/chatcompletion/apirouter.py:41-61) -> OpenAIServingChat.createchatcompletion -> createchatcompletion -> renderchatrequest (vllm/entrypoints/openai/chatcompletion/serving.py:206-280) -> OnlineRenderer.renderchat (vllm/renderers/onlinerenderer.py:95-190) -> preprocesschat (vllm/renderers/onlinerenderer.py:335-380) -> BaseRenderer.renderchatasync (vllm/renderers/base.py:1070-1105) -> HfRenderer.rendermessagesasync (vllm/renderers/hf.py:1049-1085) -> parsechatmessagesasync (vllm/entrypoints/chatutils.py:1911-1945) -> content-part parsepromptembeds and loadpromptembedsasync (vllm/entrypoints/chatutils.py:1099-1120) -> AsyncMultiModalItemTracker.resolveitems (vllm/entrypoints/chatutils.py:818-835) -> asyncio.gather of both prompt parts -> safeloadpromptembedsasync -> makeasync -> loop.runinexecutor(executor=None, ...) (vllm/utils/asyncutils.py:28-45) -> guarded torch.load -> todense().
The prompt async helper is created without an explicit executor, so it uses the event loop's default executor. This path is separate from the renderer's configurable pool. The deterministic scheduler run observed the two parts on distinct default-executor threads while leaving renderernumworkers at its default of one.
promptembeds bypasses multimodal processing, and the tracker explicitly permits it when ismultimodalmodel=False (vllm/entrypoints/chatutils.py:793-837). Consequently, the primary trigger needs neither a multimodal model nor enablemmembeds.
The source also states that async wrappers must be thread-safe (vllm/utils/asyncutils.py:28-38), while a target test acknowledges that the sparse flag is not thread-local and concurrent users can leak state (tests/renderers/testsparsetensorvalidation.py:58-61).
Exploitability Analysis
The following evidence labels separate what was demonstrated from what remains conditional:
| Label | Claim | | --- | --- | | Verified by run | PyTorch 2.11.0 rejects the identical invalid payload through the actual target loader without the race. | | Verified by run | The hash-verified frozen tracker schedules two prompt parts on distinct default-executor threads, races the flag to False, reconstructs the invalid sparse tensor, and reaches the target todense() call while the interception prevents execution. | | Traced in source | A client can supply multiple promptembeds content parts through the stock /v1/chat/completions route and the function chain above. | | Traced in source | enablepromptembeds defaults to False (vllm/config/model.py:255-260), so the operator must opt in. enablemmembeds and non-default renderer workers are not preconditions for this path. | | Traced in source | apikey defaults to None (vllm/entrypoints/openai/cliargs.py:264), and authentication middleware is installed only when a CLI or environment key is present (vllm/entrypoints/openai/apiserver.py:306-310). With a configured key, the attacker must authenticate; without one, the stock route has no API-key middleware. | | Unrun | A live HTTP/GPU server, real-world race win rate, unsafe dense conversion, process crash, memory corruption, and reliable code execution. |
The feature is documented for trusted users, which narrows intended exposure. It is not a memory-safety boundary: a user authorized to submit embedding inputs should not be able to disable a process-wide invariant for concurrent work.
The current run proves the same invalid sparse object can cross the guard and reach the historical sink. If executing that sink retains the behavior described in CVE-2025-62164 for the deployed PyTorch build, denial of service or memory corruption may follow. This is a conditional impact statement, not a reproduced outcome. Reliable RCE is not claimed.
The opt-in feature, scheduling requirement, and absence of a measured live win rate support Medium/P2 despite the serious historical sink class. No additional deployment assumptions are required for the one-request scheduler beyond stock default-executor concurrency being available.
Remediation
The immediate fix is one shared process-wide lock around every use of this process-global sparse guard. The lock must cover invariant enabling, deserialization, tensor type validation, and dense conversion:
python with sharedsparseloadlock: with torch.sparse.checksparsetensorinvariants(): tensor = torch.load(..., weightsonly=True, maplocation="cpu") validatetensortype(tensor) tensor = tensor.todense()
Every prompt, image, and audio loader that manipulates the same global flag must use the same lock. A lock only around torch.load, separate per-loader locks, or a lock omitted from the chat helper would leave overlapping save/restore sequences possible.
The stronger design is to avoid mutable process-global validation state in concurrent request code. Prefer a PyTorch per-call invariant check if one is available, or reconstruct and validate serialized embeddings inside a deliberately serialized boundary before any sparse operation.
Regression coverage should:
- Preserve the actual-target negative control using the identical malformed payload. - Force A-enter, B-enter, A-exit, B-load and assert B remains protected. - Execute the multi-part chat tracker with the event loop's default executor and renderernumworkers=1. - Cover cross-loader overlap so later prompt, image, or audio changes cannot bypass a shared fix. - Assert the global flag is restored after success and exceptions. - Reject invalid tensors before any dense conversion.
Until a fix is deployed, leaving enablepromptembeds disabled removes this stock source path.
Summary
The affected vLLM revision uses a process-global PyTorch context as the follow-up protection for CVE-2025-62164. A later chat feature causes two prompt-embedding parts from one request to run concurrently on the default executor. One context can restore the flag to False while the other is still guarded, allowing the historical malformed sparse payload class to reach the historical todense() sink. The new issue is the concurrent guard bypass and shipped trigger, not the payload or sink. Runtime validation proves the bypass and safe sink reachability on PyTorch 2.11.0; historical crash and memory-corruption effects remain conditional, and RCE was not tested or claimed.
Other sources
vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safeloadpromptembeds in vllm/renderers/embedutils.py uses torch.sparse.checksparsetensorinvariants, whose process-global save, enable, and restore state can be raced by concurrent promptembeds parts submitted to POST /v1/chat/completions through AsyncMultiModalItemTracker.resolveitems, asyncio.gather, and the default executor, allowing an invalid sparse tensor to reach tensor.todense despite the CVE-2025-62164 guard when enablepromptembeds is enabled. This issue is fixed in version 0.26.0.
— NVD
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.26.0 - Upgrade
Upgrade
vllmto a version that resolves this vulnerability.Fixed in 0.26.0 - Configuration
Until the fix is deployed, keep the default-off prompt embeddings feature disabled by setting enable_prompt_embeds to false, since the stock source path is reachable only when enable_prompt_embeds is enabled.
vllm enable_prompt_embeds = false - Compensating control
If upgrading to 0.26.0 is not immediately possible, deploy a compensating control by ensuring prompt-embedding loading work cannot overlap: use a single shared process-wide lock around every use of the process-global sparse-invariant guard that covers invariant enabling, deserialization (torch.load(weights_only=True)), tensor type validation, and dense conversion (to_dense()).
Event History
Frequently Asked Questions
What is the severity of CVE-2026-73557?
The severity of CVE-2026-73557 is rated at 65.
How do I fix CVE-2026-73557?
To fix CVE-2026-73557, update vLLM to version 0.26.0 or later.
What type of vulnerability is CVE-2026-73557?
CVE-2026-73557 is classified as a Race Condition vulnerability.
What versions of vLLM are affected by CVE-2026-73557?
CVE-2026-73557 affects vLLM versions from 0.20.2rc0 until 0.26.0.
What components are involved in CVE-2026-73557?
CVE-2026-73557 specifically involves the safe_load_prompt_embeds function in vllm/renderers/embed_utils.py.