CVE-2026-105752: vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle

Published Oct 5, 2026
·
Updated

Affected

- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.

Summary

On the GPT-OSS "Harmony" path (POST /v1/responses), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries request.cachesalt, but the tool-continuation re-submission rebuilds the engine input via tokensinput(tokenids) with no cachesalt. The continuation prefix is therefore cached in the global unsalted namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage — restoring the prompt-membership oracle that cachesalt is documented to prevent.

Silently dropping a preserved salt after the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.

This is distinct from GHSA-4qjh-9fv9-r85r (CVE-2025-46570): that advisory is the prefix-cache membership oracle for which cachesalt is the documented mitigation, and its PR-17045 fix does not close this site — the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact cachedtokensperturn counts from a different sink (the Responses serving continuation, not general TTFT timing).

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

- The drop (sink): vllm/entrypoints/openai/responses/serving.py#L712-L713 — tokenids = context.renderforcompletion() then engineinput = tokensinput(tokenids), with no cachesalt. - Correct turn-1 call for contrast: vllm/entrypoints/openai/responses/serving.py#L755 — tokensinput(prompttokenids, cachesalt=request.cachesalt). - tokensinput stores the salt only if passed: vllm/inputs/engine.py#L51-L66 (if cachesalt is not None: inputs["cachesalt"] = cachesalt). - The engine request copies only the current input's salt: vllm/v1/engine/inputprocessor.py#L380 (cachesalt=decoderinputs.get("cachesalt") → None for the continuation). - Prefix-cache hashing keys on the salt only when present: vllm/v1/core/kvcacheutils.py#L560-L561 ([request.cachesalt] if (starttokenidx == 0 and request.cachesalt) else []). - The oracle the attacker reads: vllm/entrypoints/openai/responses/serving.py#L909 (cachedtokensperturn). - The documented control being defeated: vllm/entrypoints/openai/responses/protocol.py#L235 (cachesalt field).

The tool-continuation re-submission rebuilds the engine input with no cachesalt:

python vllm/entrypoints/openai/responses/serving.py Lines 711-715 if isinstance(context, HarmonyContext): tokenids = context.renderforcompletion() engineinput = tokensinput(tokenids)

samplingparams.maxtokens = maxmodellen - len(tokenids)

Contrast with the correct turn-1 call, which does preserve the caller's salt:

python vllm/entrypoints/openai/responses/serving.py Lines 754-755 prompttokenids = renderforcompletion(messages) engineinput = tokensinput(prompttokenids, cachesalt=request.cachesalt)

tokensinput stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:

python vllm/inputs/engine.py Lines 51-66 def tokensinput( prompttokenids: list[int], , prompt: str | None = None, cachesalt: str | None = None, ) -> TokensInput: """ Construct [TokensInput][vllm.inputs.engine.TokensInput] from optional values. """ inputs = TokensInput(type="token", prompttokenids=prompttokenids)

if prompt is not None: inputs["prompt"] = prompt if cachesalt is not None: inputs["cachesalt"] = cachesalt

Impact

An authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency — the exact prompt-membership oracle cachesalt is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.

Preconditions: a GPT-OSS Harmony model on /v1/responses; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets cachesalt and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The AC:H metric reflects that guessable-history precondition.

Suggested Fix

Propagate request.cachesalt into every Harmony (and Parsable) tool-continuation re-submission — at the continuation call site call tokensinput(tokenids, cachesalt=request.cachesalt), mirroring the correct turn-1 call. Carry the salt on the HarmonyContext (thread the originating request into the context) so no continuation path can omit it:

diff vllm/entrypoints/openai/responses/serving.py if isinstance(context, HarmonyContext): tokenids = context.renderforcompletion() - engineinput = tokensinput(tokenids) + engineinput = tokensinput( + tokenids, + cachesalt=( + context.request.cachesalt + if context.request is not None + else None + ), + )

with HarmonyContext.init gaining a request: ResponsesRequest | None = None parameter (stored as self.request) that createresponses passes when constructing the context. The continuation prefix is then cached in the victim's salted namespace, mirroring turn 1.

Suggested regression test: assert cachedtokensperturn == 0 for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.

---

Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818

Other sources

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cachesalt value, placing the continuation prefix in the global unsalted cache namespace even when the caller enabled salting. On deployments with prefix caching enabled, which is the default, an authenticated tenant who can reconstruct a victim's low-entropy post-tool history can submit the same continuation and use the cachedtokensperturn count to determine whether the prefix was previously processed, defeating the intended tenant isolation of salted prefix caching. This issue is fixed in version 0.30.0.

— MITRE

Affected Software

2 affected componentsFixes available
vllm vllm<0.30.0
pip/vllm<0.30.0
0.30.0

Remediation

Recommended actions to resolve this vulnerability, in priority order.

  1. Upgrade

    Upgrade pip/vllm to a version that resolves this vulnerability.

    Fixed in 0.30.0

Event History

Oct 5, 2026
CVE Published
via MITRE·10:32 PM
Data Sourced
via MITRE·10:32 PM
DescriptionSeverityWeakness
Data Sourced
via NVD·11:17 PM
DescriptionSeverityWeakness
Oct 6, 2026
Advisory Published
via GitHub·12:02 AM
Data Sourced
via GitHub·12:02 AM
DescriptionSeverityWeaknessAffected Software

Frequently Asked Questions

1

Which deployments are exposed?

Deployments running vLLM versions before 0.30.0 with prefix caching enabled are exposed; prefix caching is enabled by default. The affected request path is Harmony tool continuations submitted to POST /v1/responses when callers use cache_salt.

2

What does an attacker need to exploit this issue?

An attacker must be an authenticated tenant and be able to reconstruct a victim's low-entropy post-tool conversation history. They can submit a matching continuation and inspect cached_tokens_per_turn to infer whether that prefix was previously processed.

3

What information can be exposed?

The issue creates a prefix-cache membership oracle. It can reveal whether a reconstructable victim continuation prefix exists in the cache, rather than exposing the cached content directly.

4

How can the issue be remediated?

Upgrade vLLM to version 0.30.0, which fixes preservation of cache_salt for Harmony tool continuations. If upgrading is not immediately possible, disabling prefix caching removes the default cache behavior required for the oracle.

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203