GHSA-7m6h-x95x-82q5: Integer Overflow
Summary An integer overflow in the actandmulkernel kernel can cause the output of one user request to be incorporated into the response of another request within the same inference batch. Under certain conditions, the last request in a batch can receive a partial or complete copy of the first user's inference result, resulting in cross-user data leakage.
Details The root cause is an integer overflow in the expression blockIdx.x 2 d at https://github.com/vllm-project/vllm/blob/ff712f6447093d07747c88680b9d006b119f5890/csrc/activationkernels.cu#L82.
As a result, the computation for one user (User A) can incorrectly consume input data from another user (User B). In particular, when 2^32 is divisible by d, the overflow can cause User A's output to contain portions of User B's inference result. In some cases, User B's response may be copied entirely into User A's response.
This constitutes a severe cross-user information disclosure vulnerability and is straightforward to trigger. PoC We reproduced the issue using meta-llama/Llama-3.2-1B-Instruct, for which d = 8192.
Using the following configuration:
Batch size: 17 Sequence length: 16384
The final response in the batch becomes an exact copy of the first response in the batch, demonstrating complete cross-user data leakage.
Impact This vulnerability enables cross-user information disclosure. An attacker can intentionally craft requests that are processed within the same inference batch as a victim's request and cause the victim's inference output to be copied into the attacker's response.
As a result, sensitive information contained in another user's model response may be exposed to an unauthorized party.
Versions
For versions prior and equal to 0.21.0, the bug is in csrc/activationkernels.cu, and for versions later than 0.21.0, the bug is in csrc/libtorchstable/activationkernels.cu.
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.27.0
Event History
Frequently Asked Questions
Which deployments are exposed to cross-user leakage?
Deployments that process multiple users' requests in the same inference batch are exposed under the affected conditions. The disclosed behavior can cause the final request in a batch to receive part or all of the first user's inference result.
What conditions are needed to trigger the issue?
The overflow occurs when 2^32 is divisible by the model dimension d. The published reproduction used meta-llama/Llama-3.2-1B-Instruct with d = 8192, a batch size of 17, and a sequence length of 16384.
Does exploitation require authentication or user interaction?
The supplied vector lists no privileges required and user interaction required. Exploitation is rated as having high attack complexity, while the advisory describes the condition as straightforward to trigger when the relevant batching and dimension conditions are met.
What is the impact if exploitation succeeds?
An attacker may receive portions of another user's inference output, and in some cases the other user's complete response. The supplied severity vector indicates high confidentiality impact, with no integrity or availability impact.