GHSA-4hhp-h66f-j5j7: SSRF
Summary
vllm/transformersutils/processors/mimov2omni.py — the multimodal processor for MiMoV2OmniForCausalLM — issues requests.get(...) directly on user-supplied image and audio URL strings and Image.open(...) on user-supplied local paths, without the SSRF / allowedlocalmediapath checks that vllm.multimodal.utils.MediaConnector was hardened with in GHSA-qh4c-xf7m-gxfc, GHSA-v359-jj2v-j536, and GHSA-pf3h-qjgv-vcpr.
This is the same bug class as those three published advisories, in a code path the patches missed. When a user passes a URL or local-file string through multimodaldata (e.g. LLM.generate(multimodaldata={"image": "http://..."})), the processor takes the unsanitized string and dispatches it without any URL-scheme allowlist, network-target allowlist, size cap, or local-path allowlist.
Details
File: vllm/transformersutils/processors/mimov2omni.py (current main)
Sink 1 — image SSRF + local-file read (fetchimage, lines 231–249):
python def fetchimage(src: Any) -> Image.Image: if isinstance(src, Image.Image): return torgb(src) if isinstance(src, bytes): return torgb(copy.deepcopy(Image.open(BytesIO(src)))) if isinstance(src, str): if src.startswith(("http://", "https://")): r = requests.get(src, timeout=30) # SSRF: no allowlist, follows redirects r.raiseforstatus() return torgb(copy.deepcopy(Image.open(BytesIO(r.content)))) if src.startswith("file://"): return torgb(Image.open(src[7:])) # arbitrary local file read if src.startswith("data:image"): ... return torgb(Image.open(src)) # fallback also opens local files raise ValueError(f"Unrecognized image source: {type(src)}")
Sink 2 — audio SSRF (around line 471):
python elif audio.startswith(("http://", "https://")): r = requests.get(audio, timeout=30) # SSRF: same pattern r.raiseforstatus() fileobj = io.BytesIO(r.content)
Reachability. fetchimage is invoked from MiMoVLProcessor.processimage:
python def processimage(self, image: ImageInput) -> torch.Tensor: kw = self.resolveimgkw(image) src = image.image if isinstance(src, (str, bytes)): src = fetchimage(src) ...
MiMoVLProcessor is wrapped by MiMoV2OmniMultiModalProcessor and registered for the MiMoV2OmniForCausalLM model architecture (vllm/modelexecutor/models/mimov2omni.py:1169). Whenever a user passes a string into multimodaldata["image"] (or ["audio"]) for this model, the unsanitized URL/path reaches the sink.
Comparison to the recent fixes. The remediation pattern adopted in the three earlier advisories was to route every external resource fetch through MediaConnector, which checks allowedlocalmediapath and applies SSRF protection before issuing the network request. chatutils.py (lines 838, 902, 924, 963, 1053, 1081) already uses self.connector.fetchimage / fetchaudio / fetchvideo. The model processor in mimov2omni.py was added later and skipped the connector — it calls requests.get and Image.open directly. Result: the public OpenAI chat-completion path is protected, but library use (LLM.generate(multimodaldata=...)), batch processing, and any other path that lets a string reach the processor receive no protection.
Impact
1. SSRF — internal-network probing / cloud-metadata theft. Standard requests.get follows redirects and accepts any URL. An attacker who controls a multimodaldata value can: - read AWS / GCP / Azure instance metadata (e.g. http://169.254.169.254/latest/meta-data/iam/security-credentials/), - probe internal services on the vLLM host (http://127.0.0.1:<port>, http://10.x.y.z), - exfiltrate via DNS / HTTP timing oracles even when the body is rejected by Image.open. 2. Arbitrary local file read via file://path (line 242) and the unguarded fallback Image.open(src) (line 248). Any file readable by the vLLM process is reachable through the model pipeline; with suitable formats this exposes /etc/passwd, ~/.aws/credentials, etc. 3. Server-side traffic generation / amplification by hammering arbitrary URLs from the vLLM host, with a 30-second timeout per request.
Suggested remediation
Replace direct requests.get and bare Image.open paths with MediaConnector.fetchimage / fetchaudioasync (or pass the inputs through MediaConnector before they reach the processor):
python vllm/transformersutils/processors/mimov2omni.py from vllm.multimodal.utils import MediaConnector
connector = MediaConnector()
def fetchimage(src): if isinstance(src, Image.Image): return torgb(src) if isinstance(src, bytes): return torgb(copy.deepcopy(Image.open(BytesIO(src)))) if isinstance(src, str): return torgb(connector.fetchimage(src)) # delegates to the hardened path raise ValueError(f"Unrecognized image source: {type(src)}")
Same change for the audio loader at line 471. This re-uses the SSRF allowlist, allowedlocalmediapath policy, and size caps that the previous patches added.
Alternative: forbid str src from reaching the processor and require all multi-modal pre-processing to go through chatutils.py / MediaConnector before hitting the model. Larger surface change, but completes the architectural fix.
Discovery
Static review on vllm@main (HEAD as of 2026-04-30) — found by triaging the file list against the three recent SSRF advisories: the mimov2omni.py processor, added after those fixes, reintroduced the same bypass class.
Reporter
Ievgen Bondarenko — sactransport2000@gmail.com — GitHub @ibondarenko1
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/vllmto a version that resolves this vulnerability.Fixed in 0.26.0 - Configuration
Route all user-supplied multi-modal image/audio inputs in `vllm/transformers_utils/processors/mimo_v2_omni.py` through `MediaConnector` so local paths are checked against `allowed_local_media_path` (and SSRF protections apply) instead of calling `requests.get(...)` and `Image.open(...)` directly.
vllm.multimodal.utils.MediaConnector allowed_local_media_path = use MediaConnector-mediated access for all local paths (enforce allowlist) - Configuration
In `vllm/transformers_utils/processors/mimo_v2_omni.py`, replace the direct `requests.get(...)` calls on user-controlled image/audio URL strings and the fallback `Image.open(src)` / `Image.open(src[7:])` local-file reads with the hardened connector methods (e.g., `_connector.fetch_image(src)` and the equivalent audio fetch via `MediaConnector`). Ensure `src`/`audio` strings cannot reach the sink without passing through `MediaConnector`.
MiMoV2OmniForCausalLM multimodal processor (vllm/transformers_utils/processors/mimo_v2_omni.py) _fetch_image / audio loader network+file handling = Use `_connector.fetch_image` / `fetch_audio_async` rather than direct `requests.get` + bare `Image.open`
Event History
Frequently Asked Questions
Which deployments are exposed in practice?
Deployments using MiMoV2OmniForCausalLM are exposed when they pass user-controlled image or audio URL strings, or local-file strings, through multi_modal_data. The affected processor dispatches those values directly rather than using the hardened MediaConnector checks.
What does an attacker need to exploit this?
An attacker needs the ability to supply a value in multi_modal_data, such as an image URL in LLM.generate(multi_modal_data={"image": "http://..."}). No user interaction is required, but the advisory’s vector specifies low privileges.
What can the flaw allow an attacker to access?
User-supplied URLs can cause server-side requests without URL-scheme or network-target allowlists, enabling SSRF. User-supplied local paths can be opened without a local-path allowlist, potentially exposing files readable by the vLLM process.
How can teams reduce exposure before applying a fix?
Do not allow untrusted users to provide URL or local-path strings through multi_modal_data for this processor. Restrict inputs to media already validated and controlled by the application rather than passing arbitrary remote URLs or filesystem paths.
How can I determine whether my application is using the affected path?
Check whether it uses MiMoV2OmniForCausalLM and sends image or audio strings through multi_modal_data. The relevant implementation is vllm/transformers_utils/processors/mimo_v2_omni.py, including its image-fetching path.