vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level media_io_kwargs field to select the GLMGA video backend and supply large values for the fps and max_frames options without a strict work ceiling. GLMGA constructs and deduplicates an attacker-sized pre-decode frame-index list, allowing a compact request and tiny valid video to consume disproportionate CPU time and memory in the shared media-loading executor. This issue is fixed in version 0.30.0.
The product does not properly control the allocation and maintenance of a limited resource.
Mitigation of resource exhaustion attacks requires that the target system either:
The first of these solutions is an issue in itself though, since it may allow attackers to prevent the use of the system by a particular valid user. If the attacker impersonates the valid user, they may be able to prevent the user from accessing the server in question.
The second solution is simply difficult to effectively institute – and even when properly done, it does not provide a full solution. It simply makes the attack require more resources on the part of the attacker.