CVE-2026-105755: vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integrity break and induced errors on `/score` and `/rerank`
Affected
- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.
Summary
On late-interaction /score and /rerank deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the caller-controlled X-Request-Id header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring — so the victim's documents are scored against the attacker's query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's X-Request-Id. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.
This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
- vllm/entrypoints/serve/engine/serving.py#L117-L124 — baserequestid() copies the public X-Request-Id header directly. - vllm/entrypoints/pooling/base/serving.py#L109 — the frontend request id is f"{self.requestidprefix}-{self.baserequestid(rawrequest)}". - vllm/entrypoints/pooling/scoring/serving.py#L211 — flashlateinteraction() (at L191) derives worker cache keys directly from that id: querykeys = [f"{ctx.requestid}-query-{i}" for i in range(nqueries)]. - vllm/v1/pool/lateinteraction.py#L30-L36 — the data-parallel routing helper pins all requests sharing a querykey to the same engine via crc32(querykey), making collisions deterministic. - vllm/v1/worker/gpu/pool/lateinteractionrunner.py#L95 — the worker stores query embeddings in a process-local cache keyed only by that string: self.querycache[querykey] = output.clone().
The caller-controlled header enters as the request id, and the flash late-interaction path derives the worker cache key directly from it:
python vllm/entrypoints/serve/engine/serving.py Lines 116-126 @staticmethod def baserequestid( rawrequest: Request | None, default: str | None = None ) -> str | None: """Pulls the request id to use from a header, if provided""" if rawrequest is not None and ( (reqid := rawrequest.headers.get("X-Request-Id")) is not None ): return reqid
return randomuuid() if default is None else default
python vllm/entrypoints/pooling/scoring/serving.py Lines 207-212 nqueries = ctx.nqueries ndocs = len(ctx.engineinputs) - nqueries queryengineinputs = ctx.engineinputs[:nqueries]
querykeys = [f"{ctx.requestid}-query-{i}" for i in range(nqueries)] queryuses = [ndocs if nqueries == 1 else 1] nqueries
The worker then stores and reads the query embedding under that string with no check that the reader owns the entry — a colliding key returns another request's cached query, or (once the use counter is exhausted) raises a cache-miss error:
python vllm/v1/worker/gpu/pool/lateinteractionrunner.py Lines 91-107 if mode == LATEINTERACTIONMODECACHEQUERY: assert queryuses is not None # output can be a view into the current step's hidden-states # buffer, so clone it before storing across scheduling steps. self.querycache[querykey] = output.clone() self.queryuses[querykey] = queryuses outputs[i] = torch.zeros((), device=output.device, dtype=torch.float32) continue
if mode == LATEINTERACTIONMODESCOREDOC: queryoutput = self.querycache.get(querykey) if queryoutput is None: raise ValueError( "late-interaction query cache miss for key " f"{querykey!r}. Ensure query requests are executed " "before their paired document requests." )
The bug is specific to the flash late-interaction path. Non-flash late-interaction scoring computes MaxSim directly from one request's in-memory outputs and does not create a cross-request worker cache key.
Impact
A network client of the standard scoring API can, on a flash late-interaction /score or /rerank deployment:
1. Corrupt another user's results — by reusing the victim's X-Request-Id, the attacker's query embedding overwrites the victim's cached entry, so the victim's documents are scored against the attacker's query (a cross-request integrity break). 2. Induce errors — depending on timing, one request consumes the shared use counter and forces the other request into a late-interaction cache-miss error.
Both consequences follow deterministically from reusing the victim's header value, because same-key work is pinned to one engine. This is reachable through normal request handling and does not depend on any trusted inter-node network.
Suggested Fix
Derive the flash late-interaction query-cache key from a server-generated, unforgeable per-request identifier (a randomuuid() namespace) rather than the caller-supplied X-Request-Id, and thread that key through the PoolingServeContext to the doc-scoring pass so both passes reuse the same key and a caller cannot address another request's cache entry.
In vllm/entrypoints/pooling/scoring/serving.py, the encode-queries pass mints a fresh namespace and stashes the keys on the context:
python - querykeys = [f"{ctx.requestid}-query-{i}" for i in range(nqueries)] + querynamespace = randomuuid() + querykeys = [ + f"late-interaction-{querynamespace}-query-{i}" for i in range(nqueries) + ] + ctx.lateinteractionquerykeys = querykeys
and the encode-docs pass reads those stored keys instead of re-deriving them from ctx.requestid:
python - querykeys = [f"{ctx.requestid}-query-{i}" for i in range(nqueries)] + querykeys = ctx.lateinteractionquerykeys + if querykeys is None: + raise RuntimeError("Late-interaction query keys were not initialized.")
This requires adding the lateinteractionquerykeys: list[str] | None = None field to PoolingServeContext (vllm/entrypoints/pooling/typing.py). Because the namespace is a server-generated UUID, colliding X-Request-Id values no longer produce a shared cache key; a regression test asserting exactly that (colliding request ids yield distinct query-cache keys) accompanies the change.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
---
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51445
Other sources
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's querykey value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim's identifier can overwrite the cached query embedding so the victim's documents are scored against the attacker's query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.30.0
Event History
Frequently Asked Questions
Which requests are exposed?
Deployments using flash late-interaction scoring through the /score or /rerank endpoints are affected before version 0.30.0. The issue involves the X-Request-Id header used to derive a worker query cache key.
What must an attacker do to interfere with another request?
The attacker needs to send a concurrent request that reuses the victim request's X-Request-Id value. This can overwrite the cached query embedding, causing the victim's documents to be scored using the attacker's query.
What are the operational effects besides incorrect scoring?
Shared use counters can trigger a late-interaction cache-miss error. This can induce errors for affected /score and /rerank requests.
What version resolves the issue?
Upgrade vLLM to version 0.30.0, which fixes the issue.