Open-Source Security Intelligence

Know every vulnerability
before it knows you.

DevGuard continuously monitors your dependencies and alerts you when CVEs like this one affect your stack — with real-time threat intelligence built for developers.

Search

GHSA-hcwq-8wjf-3gcr

MediumCVSS 6.5 / 10
Published Sep 16, 2026·Last modified Sep 16, 2026
Affected Components(1)
PyPI logovllm
< 0.24.0
Description

Summary

The audio decode-duration guard (max_duration_s, env VLLM_MAX_AUDIO_DECODE_DURATION_S, default 600s) that protects against audio decompression-bomb DoS is wired into only the speech-to-text path (/v1/audio/transcriptions). The chat audio path (/v1/chat/completions, input_audio content parts) calls the same decoder with no limit, so an unauthenticated client can submit a few-KB compressed audio file that expands to multiple GB of float32 PCM at decode time, OOM-killing the worker. This is a distinct sibling of CVE-2026-5497 (video frame-count bomb, VideoMediaIO.load_base64) and GHSA-pq5c-rjhq-qp7p (image) in the same media subsystem.

Verified against main at HEAD d78650c (2026-06-16); applicable to the latest release v0.23.0.

Details

The guard rejects long audio during decode (before allocation), implemented in vllm/multimodal/media/audio.py:

  • load_audio_pyav — metadata reject (~82-98) and live sample-count reject (~129-136)
  • load_audio_soundfile — frames reject (~165-174)

All are gated on if max_duration_s is not None.

It is passed in exactly one place — the transcription serving layer:

# .../speech_to_text/base/serving.py:~170-174
load_audio(buf, sr=..., max_duration_s=self.max_audio_decode_duration_s)
#   self.max_audio_decode_duration_s = envs.VLLM_MAX_AUDIO_DECODE_DURATION_S  (default 600)

The chat path never threads it:

# vllm/multimodal/media/audio.py:237-238
def load_bytes(self, data: bytes) -> tuple[npt.NDArray, float]:
    return load_audio(BytesIO(data), sr=None)   # no max_duration_s -> every guard above is skipped

Unauthenticated reachability chain (chat): parse_input_audio (chat_utils.py) -> parse_audio -> connector.fetch_audio -> AudioMediaIO._load_data_url -> load_base64 -> load_bytes -> load_audio(..., sr=None). The connector never passes max_duration_s, and inline data: URLs need no HTTP fetch (so VLLM_AUDIO_FETCH_TIMEOUT does not bound them). The OpenAI-compatible server has no auth by default (auth only when --api-key / VLLM_API_KEY is set).

Impact

Unauthenticated remote denial of service (availability) via memory amplification on a default-no-auth endpoint, on any deployment serving an audio-capable model. Same class and impact as the sibling CVE-2026-5497 (video). CWE-770 / CWE-409.

Fix

A fix was introduced in this MR: https://github.com/vllm-project/vllm/pull/45908

Upload your SBOM

Upload your own SBOM in CycloneDX 1.6 or higher (JSON) directly here to check your vulnerabilities.

Risk Scores
Base Score
6.5

The vulnerability can be exploited over the network without needing physical access. It is easy for an attacker to exploit this vulnerability. An attacker needs basic access or low-level privileges. No user interaction is needed for the attacker to exploit this vulnerability. The impact is confined to the system where the vulnerability exists. There is a high impact on the availability of the system.

Threat Intelligence
6.0

Exploitation attempts have been detected. Elevated vigilance and prompt remediation are advised.

EPSS
0.69%

The exploit probability is very low. The vulnerability is unlikely to be exploited in the next 30 days.

Exploit
Not available

We did not find any exploit available. Neither in GitHub repositories nor in the Exploit-Database.

Browse More

Scan your project

Continuously monitor your dependencies and get alerted when vulnerabilities like this one affect your stack.

Checkout DevGuard