vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the Rust frontends track_http_metrics middleware records the raw HTTP method token as a Prometheus label for requests reaching registered routes. An unauthenticated attacker can send unique arbitrary method tokens to unguarded routes such as /tokenize, causing Prometheuss Family::get_or_create function to permanently create counter and histogram label sets. Those label sets increase process memory usage and enlarge the /metrics response until the service or monitoring path is exhausted. This issue is fixed in version 0.30.0.
The product does not properly control the allocation and maintenance of a limited resource.
Mitigation of resource exhaustion attacks requires that the target system either:
The first of these solutions is an issue in itself though, since it may allow attackers to prevent the use of the system by a particular valid user. If the attacker impersonates the valid user, they may be able to prevent the user from accessing the server in question.
The second solution is simply difficult to effectively institute – and even when properly done, it does not provide a full solution. It simply makes the attack require more resources on the part of the attacker.