Know every vulnerabilitybefore it knows you.
DevGuard continuously monitors your dependencies and alerts you when CVEs like this one affect your stack — with real-time threat intelligence built for developers.
GHSA-2823-qmq8-rwvj
Affected
- Ecosystem / package: pip /
vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit
752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the loosecache_saltvalidator and unguarded scheduling-path lookup reach.
Summary
vLLM's OpenAI-compatible request models (Completions, Chat Completions, Responses) accept a client-supplied cache_salt field and validate it only as "must be a non-empty string" — no character or length restrictions. On a deployment with the built-in LMCache-MP KV connector enabled, that value is stored verbatim on the request tracker and forwarded unguarded as a keyword argument into the scheduler's per-step cache lookup. The downstream LMCache library applies a stricter check in IPCCacheServerKey.__post_init__, which raises ValueError for any cache_salt containing @, /, \, or NUL (or longer than 128 characters).
Neither the LMCache-MP connector lookup call site nor Scheduler.schedule() wraps that call in a request-scoped try/except, so the ValueError propagates uncaught into EngineCore's top-level handler, which treats any uncaught exception as fatal and kills the whole engine process. A single publicly reachable request with, for example, cache_salt="/" therefore takes down the engine for all concurrent users. vLLM's boundary validator is looser than the downstream consumer's, and the gap is never converted into a request-scoped failure on the scheduling path.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
- The three
check_cache_salt_supportvalidators require only a non-empty string — no character or length bound:vllm/entrypoints/openai/completion/protocol.py#L502-L508(field at#L172),vllm/entrypoints/openai/chat_completion/protocol.py#L913-L919(field at#L425), andvllm/entrypoints/openai/responses/protocol.py#L459-L465(field at#L235). The loose test itself is atcompletion #L503-L505,chat_completion #L914-L916,responses #L460-L462. - Two further request models accept
cache_saltwith the same or weaker checking, and should be hardened at the same time:vllm/entrypoints/pooling/base/protocol.py#L74-L85carries the identical non-empty-string-only validator (field at#L59), and the token-in-token-out scale-outGenerateRequestexposes the field with nocache_saltvalidator at all:vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L110. - The LMCache-MP request tracker stores the salt verbatim:
vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L189(LMCacheMPRequestTracker), assignment at#L223. - The lookup call forwards
cache_saltwith no surroundingtry:vllm/distributed/kv_transfer/kv_connector/v1/lmcache_mp_connector.py#L733-L774(get_num_new_matched_tokens→maybe_submit_lookup_request(...)at#L770). - The scheduler invokes the connector unguarded:
vllm/v1/core/sched/scheduler.py#L739(self.connector.get_num_new_matched_tokens(...), insideschedule()at #L396). - The generic top-level handler treats the exception as fatal:
vllm/v1/engine/core.py#L1229-L1235— theexcept Exceptioninrun_engine_core(#L1154) logsEngineCore encountered a fatal error., calls_send_engine_dead(), and re-raises. - Downstream strict validator (external LMCache library, not vLLM):
IPCCacheServerKey.__post_init__inlmcache/v1/multiprocess/custom_types.pyrejects@ / \NUL and >128-char salts (_SALT_FORBIDDEN_CHARS = frozenset("@/\\\x00")) by raisingValueError.
The three OpenAI check_cache_salt_support validators are identical; the completion one is representative — a non-empty-string test with no character or length bound:
# vllm/entrypoints/openai/completion/protocol.py Lines 503-509
if data.get("cache_salt") is not None and (
not isinstance(data["cache_salt"], str) or not data["cache_salt"]
):
raise VLLMValidationError(
"Parameter 'cache_salt' must be a non-empty string if provided.",
parameter="cache_salt",
)
On the scheduling path the connector is invoked with no surrounding try — a ValueError from the downstream salt check propagates straight out of schedule():
# vllm/v1/core/sched/scheduler.py Lines 736-742
# Get externally-cached tokens if using a KVConnector.
if self.connector is not None:
ext_tokens, load_kv_async = (
self.connector.get_num_new_matched_tokens(
request, num_new_local_computed_tokens
)
)
run_engine_core's generic handler — the next except up the stack — treats that as fatal, marks the engine dead, and re-raises:
# vllm/v1/engine/core.py Lines 1229-1235
except Exception as e:
if engine_core is None:
logger.exception("EngineCore failed to start.")
else:
logger.exception("EngineCore encountered a fatal error.")
engine_core._send_engine_dead()
raise e
Impact
Availability only. cache_salt is an attacker-controlled, publicly reachable request field that vLLM validates too loosely. A value such as "/" passes vLLM's check, reaches the stricter downstream validator, and its ValueError is never converted into a request-scoped failure — instead it kills the EngineCore process, a denial of service for every concurrent request on that server (HTTP failures, then /health failing).
Applicability: the built-in LMCache-MP KV connector must be enabled (lmcache >= 0.4.4), which is itself an opt-in KV-connector boundary. On such deployments no other special configuration is required, and the crash is a resource-availability failure rather than expected behavior of the opt-in feature.
Suggested Fix
Two independent fixes:
- Tighten admission — add one shared
validate_cache_salt()helper (for example invllm/entrypoints/openai/engine/protocol.py) that matches or exceeds the downstreamIPCCacheServerKeyrules — reject@,/,\, NUL, and >128-character salts at the HTTP boundary with a 4xx — and route every request model that exposescache_saltthrough it: the threecheck_cache_salt_supportvalidators above, the pooling base request, and the token-in-token-outGenerateRequest(which has no validator today). - Defense in depth — wrap the LMCache-MP lookup call reached from
Scheduler.schedule()so a downstream validatorValueErrorbecomes a request-scoped failure instead of an EngineCore-fatal exception. This is the fix that also covers future divergence between vLLM's and LMCache's salt rules; the right failure semantics (fail the request vs. fall back to a cold lookup) is a maintainer design call.
The core of fix 1 is a single shared helper; each check_cache_salt_support body then becomes validate_cache_salt(data.get("cache_salt")), and GenerateRequest gains an equivalent mode="before" validator:
# vllm/entrypoints/openai/engine/protocol.py — new shared helper
_CACHE_SALT_FORBIDDEN_CHARS = frozenset("@/\\\x00")
_MAX_CACHE_SALT_LENGTH = 128
def validate_cache_salt(cache_salt: object) -> None:
"""Validate cache salts before they reach downstream cache backends."""
if cache_salt is None:
return
if not isinstance(cache_salt, str) or not cache_salt:
raise VLLMValidationError(
"Parameter 'cache_salt' must be a non-empty string if provided.",
parameter="cache_salt",
)
if len(cache_salt) > _MAX_CACHE_SALT_LENGTH or any(
char in _CACHE_SALT_FORBIDDEN_CHARS for char in cache_salt
):
raise VLLMValidationError(
"Parameter 'cache_salt' must be at most 128 characters and must "
"not contain '@', '/', '\\\\', or NUL.",
parameter="cache_salt",
)
This distinguishes the finding from GHSA-6qc9-v4r8-22xg: that advisory fixed only the guided_json/xgrammar trigger of the same schedule()-into-run_engine_core fatal-handler family, so the cache_salt admission gap and the schedule()-level defense-in-depth (fix 2) survive its published fix. A patch implementing fix 1 across all five request models, with a regression test covering the rejected ("/", 129-char) and accepted salt shapes, applies to v0.25.1 with line offsets and no fuzz.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51444
Upload your own SBOM in CycloneDX 1.6 or higher (JSON) directly here to check your vulnerabilities.
Drag and drop some file here, or click to select
The vulnerability can be exploited over the network without needing physical access. It is easy for an attacker to exploit this vulnerability. An attacker needs basic access or low-level privileges. No user interaction is needed for the attacker to exploit this vulnerability. The impact is confined to the system where the vulnerability exists. There is a high impact on the availability of the system.
Exploitation attempts have been detected. Elevated vigilance and prompt remediation are advised.
The exploit probability is very low. The vulnerability is unlikely to be exploited in the next 30 days.
We did not find any exploit available. Neither in GitHub repositories nor in the Exploit-Database.
Browse More
Continuously monitor your dependencies and get alerted when vulnerabilities like this one affect your stack.
Checkout DevGuard