CVE-2026-105757: vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
Affected
- Ecosystem / package: pip / vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back each path reaches.
Summary
Three structured-output request paths let an ordinary request reach a condition that raises an uncaught, engine-fatal exception instead of a per-request validation error. The failure is not confined to the request that caused it: it escapes into EngineCore's busy loop and triggers a fatal sendenginedead(), so one malformed structured-output request denies service to all concurrent and subsequent tenants of that engine. The requests are ordinary API calls; the client does not need any privileged configuration beyond the (often default) structured-output feature.
The shared root cause is the absence of a per-request exception boundary around structured-output grammar/token handling: a value that should fail as a request-scoped validation error instead propagates as an uncaught exception (or bypasses frontend validation entirely) and reaches the scheduler/engine loop, which treats the failure as fatal.
These sites are distinct from the published fixes for GHSA-6qc9-v4r8-22xg and GHSA-8wr5-jm2h-8r4f: GHSA-6qc9 (PR #17623) added frontend json-schema/regex/type validation but the completion check on v0.25.1 still catches only TimeoutError, so Site 1's valid duplicate-root EBNF via the latched xgrammar backend still re-raises and kills EngineCore; GHSA-8wr5 (PR #44744) fixes recovered-token state in the eagle spec-decode path and explicitly does not treat trailing -1 padding as the fault, so Site 2's -1 reaching backendguidance.validatetokens() via ngramgpu survives it.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1). The three sites are:
Site 1 — latched-backend compile exception escapes. StructuredOutputManager permanently latches one backend on the first structured-output request and always compiles through it, ignoring the per-request auto backend selection. A grammar that auto accepts only via a fallback backend (e.g. a duplicate-root EBNF) then makes the latched xgrammar compiler raise; the exception is stored in the grammar Future, and the completion check catches only TimeoutError, so Future.result() re-raises during scheduler promotion and kills EngineCore.
- Backend defaults to "auto": vllm/config/structuredoutputs.py#L21. - The auto validator tries xgrammar and silently falls back, recording the resolved backend on the request: vllm/samplingparams.py#L1004-L1035 (fields at #L85-L88). - The manager latches one backend and always compiles through it: vllm/v1/structuredoutput/init.py#L127-L159 and creategrammar() at #L173-L184. - The xgrammar sink calls the native compiler unguarded: vllm/v1/structuredoutput/backendxgrammar.py#L78-L110 (grammar path at #L90). - The completion check catches only TimeoutError: vllm/v1/structuredoutput/request.py#L48-L63 (result(timeout=0.0001) at #L55, except TimeoutError at #L57). - Scheduler promotion dereferences the grammar with no exception boundary: vllm/v1/core/sched/scheduler.py#L2441-L2460 (reached from schedule() at #L396). - The uncaught exception is treated as fatal: vllm/v1/engine/core.py#L1229-L1234 (sendenginedead() at #L1470).
The escaping-exception sink: the completion check only handles TimeoutError, so any compile exception stored in the Future re-raises out of result() during scheduler promotion.
python vllm/v1/structuredoutput/request.py Lines 48-59 def checkgrammarcompletion(self) -> bool: # NOTE: We have to lazy import to gate circular imports from vllm.v1.request import RequestStatus
if isinstance(self.grammar, Future): try: # We will check whether the future is ready within 100 us self.grammar = self.grammar.result(timeout=0.0001) self.status = RequestStatus.WAITING except TimeoutError: return False return True
Site 2 — guidance treats ngramgpu -1 padding as a token id. The ngramgpu proposer pads a fixed-width draft-token row with -1 and separately records numvaliddrafttokens, but the scheduler forwards the untrimmed padded row into GuidanceGrammar.validatetokens(), which passes -1 to the native llguidance matcher; the matcher cannot convert a negative value to its unsigned token type and raises an OverflowError, uncaught and fatal.
- The proposer fills with -1 and computes numvaliddrafttokens: vllm/v1/specdecode/ngramproposergpu.py#L189-L209. - getdrafttokenidscpu() materializes the full padded row, not trimmed to the valid count: vllm/v1/worker/gpumodelrunner.py#L4835-L4849. - Two scheduler call sites pass the padded row into validatetokens(...): vllm/v1/core/sched/scheduler.py#L1967 and #L1997. - The sink has no negative-value filter: vllm/v1/structuredoutput/backendguidance.py#L181-L192 (validatetokens(tokens) at #L192). - The uncaught OverflowError is treated as fatal: vllm/v1/engine/core.py#L1229-L1234.
The sink passes the padded row straight to the native matcher with no lower-bound check on token ids:
python vllm/v1/structuredoutput/backendguidance.py Lines 181-196 def validatetokens(self, tokens: list[int]) -> list[int]: """Checks if the list of tokens are accepted by the parser in sequence. Will not advance the parser.
Returns the prefix list of tokens that are accepted by the parser. """ if len(tokens) == 0: return [] if self.llmatcher.isstopped(): return []
numtokens = self.llmatcher.validatetokens(tokens)
self.checkerror()
return tokens[:numtokens]
Site 3 — Rust frontend accepts an empty structured-output value the Python frontend rejects. The opt-in Rust frontend admits an empty structuredoutputs.json / structuredoutputs.grammar string that the Python frontend rejects; the empty value passes through the wire-params conversion without validation, reaches the engine, and marks EngineCore dead.
- impl TryFrom<WireStructuredOutputsParams> for StructuredOutputsParams: rust/src/engine-core-client/src/protocol/structuredoutputs.rs#L135-L163 — tryfrom at L138 maps raw.json to a Json constraint (L159) and raw.grammar to a Grammar constraint (L162) with no non-empty check, so an empty value is forwarded unchanged. WireStructuredOutputsParams is defined at #L115.
The conversion maps each field to a constraint with no non-empty check, so an empty json/grammar string is forwarded unchanged:
rust // rust/src/engine-core-client/src/protocol/structuredoutputs.rs Lines 135-162 impl TryFrom<WireStructuredOutputsParams> for StructuredOutputsParams { type Error = Error;
fn tryfrom(raw: WireStructuredOutputsParams) -> Result<Self> { use StructuredOutputConstraint::;
let mut constraint = None;
macrorules! insertconstraint { ($name:literal, $value:expr) => { if let Some(value) = $value { if let Some((existing, )) = constraint { return Err(Error::InvalidStructuredOutputsParams { message: format!( "multiple structured output constraints specified: {existing}, {}", $name ), }); } constraint = Some(($name, value)); } }; }
insertconstraint!("json", raw.json.map(Json)); insertconstraint!("regex", raw.regex.map(Regex)); insertconstraint!("choice", raw.choice.map(Choice)); insertconstraint!("grammar", raw.grammar.map(Grammar));
Impact
Any client able to send an ordinary structured-output request (two requests for Site 1) can terminate the shared EngineCore, denying service to all tenants of that engine instance. Availability only; no code execution, memory corruption, or data disclosure.
- Site 1 affects the default "auto" structured-output backend; no speculative decoding or non-default configuration is required. - Site 2 requires a deployment using the guidance backend together with ngramgpu speculative decoding; the -1 sentinel is produced by the proposer itself, so an ordinary constrained request suffices — the client does not craft the negative token. - Site 3 requires the opt-in Rust frontend; a single request with an empty constraint value is enough. API-key middleware narrows the attacker from unauthenticated to authenticated but does not restore the missing validation.
Suggested Fix
Convert structured-output failures into request-scoped errors at the boundary. Each site is a distinct fix shape; the core hunk for each is below.
Site 1 — resolve the backend per request instead of latching one process-wide, so a grammar is always compiled with the backend that auto actually selected for it (and a compile failure is that request's error, not the engine's):
python vllm/v1/structuredoutput/init.py — grammarinit(), replacing the latched-backend path backend = self.getorcreatebackend(request.structuredoutputrequest.backend) if self.useasyncgrammarcompilation: grammar = self.executor.submit(self.creategrammar, request, backend) else: grammar = self.creategrammar(request, backend)
Site 2 — drop negative sentinels before handing the row to the native matcher, inside validatetokens():
python vllm/v1/structuredoutput/backendguidance.py — validatetokens(), before the llmatcher call for i, token in enumerate(tokens): if token < 0: tokens = tokens[:i] break if len(tokens) == 0: return []
Site 3 — reject empty json/grammar strings in the Rust wire-params conversion, matching the Python frontend, before the value reaches the engine:
rust // rust/src/engine-core-client/src/protocol/structuredoutputs.rs — tryfrom(), before constraint mapping if matches!(&raw.json, Some(Value::String(value)) if value.trim().isempty()) { return Err(Error::InvalidStructuredOutputsParams { message: "structuredoutputs.json cannot be an empty string".tostring(), }); } if matches!(&raw.grammar, Some(value) if value.trim().isempty()) { return Err(Error::InvalidStructuredOutputsParams { message: "structuredoutputs.grammar cannot be an empty string".tostring(), }); }
The three sites share one root cause and the same fix shape (bound the failure to the request), so they are filed as a single advisory; happy to split into per-component advisories if the maintainers prefer. Each fix carries a regression test.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
These vulnerabilities were discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
---
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51450
Other sources
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngramgpu speculative-decoding mode can pass a negative token to guidance validation, and the Rust frontend can admit empty structured-output values that the Python frontend rejects, allowing ordinary constrained-generation requests to terminate the shared engine. This issue is fixed in version 0.30.0.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.30.0
Event History
Frequently Asked Questions
Who can trigger the denial of service?
An attacker needs low-privileged access to submit constrained-generation or structured-output requests to a vLLM service. The affected failures are request-driven and can terminate the shared EngineCore.
Which deployments are affected?
vLLM versions prior to 0.30.0 are affected. Exposure depends on accepting structured-output requests; the described trigger paths include backend mismatches, ngram_gpu speculative decoding padding, and the Rust frontend accepting empty structured-output values.
What is the practical impact of a successful exploit?
A malformed or incompatible structured-output request can cause an exception to reach the EngineCore fatal-error path, terminating the shared engine. The provided impact assessment indicates availability impact only, with no confidentiality or integrity impact.
What should teams do if they cannot immediately upgrade?
Restrict structured-output and constrained-generation request submission to trusted users, and avoid the described risky configurations or inputs where operationally possible. In particular, limit use of ngram_gpu speculative decoding with guidance validation and do not accept empty structured-output values through the Rust frontend.
How can operators determine whether they are affected?
Check whether the deployed vLLM version is earlier than 0.30.0 and whether the service processes structured-output requests. Engine termination associated with grammar-compilation exceptions, guidance validation involving a negative token, or empty structured-output values accepted by the Rust frontend is consistent with the described issue.