When I call the api '/v1/chat/completions' of API Server, it response incomplete results, but vllm's api response complete results
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
When I call the api '/v1/chat/completions' of API Server to access vllm_worker server , it response incomplete results, but vllm's api response complete results and model_work server response complete results
## env
- fastchat 0.2.36
- vllm 0.3.1
- python3.9
- cuda 12.2
- ubuntu20.04
- qwen1.5-72B-chat
- A800*4
## Start command
python3.9 -m fastchat.serve.vllm_worker --model-path Qwen/Qwen1.5-72B-Chat --model-names Qwen1.5-72b-chat --controller-address http://70.182.56.16:21001 --worker-address http://70.182.56.16:21004 --host 0.0.0.0 --port 21002 --tensor-parallel-size 4 --gpu-memory-utilization 0.98
## Detail
When calling the api '/v1/chat/completions' of API Server, it response incomplete results "Hello! How can I assis"
```
curl --location --request POST 'http://localhost:8000/v1/chat/completions' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "Qwen1.5-72b-chat",
"messages": [
{
"role": "user",
"content": "hello"
}
]
}'
```
```json
{
"id": "chatcmpl-BpWqkwhWbXyXbDYM2kqgdp",
"object": "chat.completion",
"created": 1708576975,
"model": "Qwen1.5-72b-chat",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist "
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"total_tokens": 30,
"completion_tokens": 10
}
}
```
Directly calling the worker_generate_stream API, the output is out of order, and the complete output is not the last one.
The last output is incomplete results
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assis".
The third to last output is complete results
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?\n",
```
curl --location --request POST 'http://70.182.56.16:21004/worker_generate_stream' \
--header 'Content-Type: application/json' \
--data-raw '{
"prompt": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\n"
}'
```
```json
{
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 1,
"total_tokens": 20
},
"cumulative_logprob": [-0.5442918539047241],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello!",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 2,
"total_tokens": 21
},
"cumulative_logprob": [-0.6264679655432701],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 3,
"total_tokens": 22
},
"cumulative_logprob": [-0.63395517738536],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 4,
"total_tokens": 23
},
"cumulative_logprob": [-0.7231398862786591],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 5,
"total_tokens": 24
},
"cumulative_logprob": [-0.7231670656274218],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 6,
"total_tokens": 25
},
"cumulative_logprob": [-1.6991723814080615],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 7,
"total_tokens": 26
},
"cumulative_logprob": [-1.6991882361180615],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 8,
"total_tokens": 27
},
"cumulative_logprob": [-1.710257318773074],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 9,
"total_tokens": 28
},
"cumulative_logprob": [-1.7102782993879373],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 10,
"total_tokens": 29
},
"cumulative_logprob": [-1.8846533119049127],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?\n",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 11,
"total_tokens": 30
},
"cumulative_logprob": [-1.8846546232061883],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assis",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 12,
"total_tokens": 31
},
"cumulative_logprob": [-2.045250159438183],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assis",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 12,
"total_tokens": 31
},
"cumulative_logprob": [-2.045250159438183],
"finish_reason": "stop"
}
```
**When using vllm's OpenAI-compatible API service, it response complete results.**
```
python3.9", "-m", "vllm.entrypoints.openai.api_server", "--model", "/Qwen/Qwen1.5-72B-Chat", "--port", "21002", "--gpu-memory-utilization", "0.98", "--tensor-parallel-size", "4
{
"id": "cmpl-a10de9e3ba27482296eecc572b5623d6",
"object": "chat.completion",
"created": 443291,
"model": "/Qwen/Qwen1.5-72B-Chat",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist you today?\n"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 9,
"total_tokens": 21,
"completion_tokens": 12
}
}
```
**When using model_work, it response also complete results.**
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start at the API Server's /v1/chat/completions path and the vllm_worker /worker_generate_stream endpoint, then compare their handling of streamed results with vllm.entrypoints.openai.api_server. Reproduce using the Qwen1.5-72B-Chat command and the supplied curl requests. Done means the API returns the complete assistant response rather than the final truncated stream chunk, with ordering and finish_reason preserved.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- api, backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100