intel / intel/llm-scaler

VLLM seems be hang

Open
#391 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
529
Forks
80
Avg merge
9h 7m
Merged PRs (30d)
38

Description

The VLLM log trace consistently displays **Running: 1 reqs, Waiting: 14 reqs** for over an hour, despite no active input prompts being sent to the model, the vllm seems be hang and what the root cause is potentially?

```
(APIServer pid=851) INFO 04-28 22:54:32 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 5.4 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 76.5%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:59146 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:59146 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 22:54:42 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 5.4 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 76.7%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:59146 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:59146 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 22:54:52 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 5.5 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 76.8%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:59146 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:59146 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 22:55:02 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 5.4 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 77.0%, Prefix cache hit rate: 92.0%

(APIServer pid=851) INFO 04-28 23:12:42 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 89.2%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:12:52 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 89.4%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:13:02 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 89.5%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:13:12 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 89.5%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:13:22 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 89.7%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:13:32 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 89.8%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:13:42 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 90.0%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:13:52 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.6 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 90.0%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:14:02 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 90.1%, Prefix cache hit rate: 92.0%
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO: 10.109.20.86:56202 - "GET /metrics HTTP/1.1" 200 OK
(APIServer pid=851) INFO 04-28 23:14:12 [loggers.py:257] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 4.7 tokens/s, Running: 1 reqs, Waiting: 14 reqs, GPU KV cache usage: 90.3%, Prefix cache hit rate: 92.0%
```

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the reported state from the APIServer logs, with no active prompts, while monitoring the repeated GET /metrics requests and the Running/Waiting counters. Trace why one request remains running and 14 remain queued as GPU KV cache usage rises; done means identifying a reproducible root cause and documenting the relevant evidence.

Written by the indexing model from the issue text.

Assessment

Domain
ai, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.