Caching Issue in vLLM's Multimodal Rendering Affects Service Availability
CVE-2026-105753
6.5MEDIUM
What is CVE-2026-105753?
vLLM, a large language model inference and serving engine, contains a vulnerability in its default mirrored multimodal LRU cache prior to version 0.28.0. The issue arises during multimodal rendering where a media hash can be committed to the frontend sender cache before receiving engine admission. If the request is rejected, the payload is never received by the engine receiver cache. Consequently, subsequent requests that reuse the same media hash may result in a failure to send the expected payload, triggering an assertion error in the receiver cache and leading to service availability issues. This vulnerability has been addressed in version 0.28.0.
Affected Version(s)
vllm < 0.28.0
