Caching Issue in vLLM's Multimodal Rendering Affects Service Availability
CVE-2026-105753

6.5MEDIUM

Key Information:

Status
Vendor
CVE Published:
5 October 2026

What is CVE-2026-105753?

vLLM, a large language model inference and serving engine, contains a vulnerability in its default mirrored multimodal LRU cache prior to version 0.28.0. The issue arises during multimodal rendering where a media hash can be committed to the frontend sender cache before receiving engine admission. If the request is rejected, the payload is never received by the engine receiver cache. Consequently, subsequent requests that reuse the same media hash may result in a failure to send the expected payload, triggering an assertion error in the receiver cache and leading to service availability issues. This vulnerability has been addressed in version 0.28.0.

Affected Version(s)

vllm < 0.28.0

References

CVSS V3.1

Score:
6.5
Severity:
MEDIUM
Confidentiality:
None
Integrity:
None
Availability:
None
Attack Vector:
Network
Attack Complexity:
Low
Privileges Required:
Low
User Interaction:
None
Scope:
Unchanged

Timeline

  • Vulnerability published

  • Vulnerability Reserved

.