Vulnerability in vLLM Inference Engine Affects Large Language Model Deployments
CVE-2026-105754
6.5MEDIUM
What is CVE-2026-105754?
The vLLM inference and serving engine for large language models contains a security vulnerability that allows attacker-controlled data inputs in the /inference/v1/generate endpoint, which can lead to various security issues including resource exhaustion, cache poisoning, and unintended alterations of transport semantics. Attackers can craft malicious inputs that exploit this flaw, especially targeting cache states and encoder behaviors. This significant issue has been addressed in vLLM version 0.30.0.
Affected Version(s)
vllm < 0.30.0
