lm-sys / lm-sys/FastChat

When I call the api '/v1/chat/completions' of API Server, it response incomplete results, but vllm's api response complete results

Open
#3,074 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

When I call the api '/v1/chat/completions' of API Server to access vllm_worker server , it response incomplete results, but vllm's api response complete results and model_work server response complete results

## env
- fastchat 0.2.36
- vllm 0.3.1
- python3.9
- cuda 12.2
- ubuntu20.04
- qwen1.5-72B-chat
- A800*4
## Start command
python3.9 -m fastchat.serve.vllm_worker --model-path Qwen/Qwen1.5-72B-Chat --model-names Qwen1.5-72b-chat --controller-address http://70.182.56.16:21001 --worker-address http://70.182.56.16:21004 --host 0.0.0.0 --port 21002 --tensor-parallel-size 4 --gpu-memory-utilization 0.98

## Detail

When calling the api '/v1/chat/completions' of API Server, it response incomplete results "Hello! How can I assis"
```
curl --location --request POST 'http://localhost:8000/v1/chat/completions' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "Qwen1.5-72b-chat",
"messages": [
{
"role": "user",
"content": "hello"
}
]
}'
```

```json
{
"id": "chatcmpl-BpWqkwhWbXyXbDYM2kqgdp",
"object": "chat.completion",
"created": 1708576975,
"model": "Qwen1.5-72b-chat",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist "
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"total_tokens": 30,
"completion_tokens": 10
}
}
```
Directly calling the worker_generate_stream API, the output is out of order, and the complete output is not the last one.

The last output is incomplete results
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assis".

The third to last output is complete results
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?\n",

```
curl --location --request POST 'http://70.182.56.16:21004/worker_generate_stream' \
--header 'Content-Type: application/json' \
--data-raw '{
"prompt": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\n"
}'
```

```json
{
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 1,
"total_tokens": 20
},
"cumulative_logprob": [-0.5442918539047241],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello!",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 2,
"total_tokens": 21
},
"cumulative_logprob": [-0.6264679655432701],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 3,
"total_tokens": 22
},
"cumulative_logprob": [-0.63395517738536],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 4,
"total_tokens": 23
},
"cumulative_logprob": [-0.7231398862786591],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 5,
"total_tokens": 24
},
"cumulative_logprob": [-0.7231670656274218],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 6,
"total_tokens": 25
},
"cumulative_logprob": [-1.6991723814080615],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 7,
"total_tokens": 26
},
"cumulative_logprob": [-1.6991882361180615],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 8,
"total_tokens": 27
},
"cumulative_logprob": [-1.710257318773074],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 9,
"total_tokens": 28
},
"cumulative_logprob": [-1.7102782993879373],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 10,
"total_tokens": 29
},
"cumulative_logprob": [-1.8846533119049127],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assist you today?\n",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 11,
"total_tokens": 30
},
"cumulative_logprob": [-1.8846546232061883],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assis",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 12,
"total_tokens": 31
},
"cumulative_logprob": [-2.045250159438183],
"finish_reason": null
}, {
"text": "<|im_start|>system\nyou are a helpful assistant<|im_end|>\n<|im_start|>user\nhello<|im_end|>\n<|im_start|>assistant\nHello! How can I assis",
"error_code": 0,
"usage": {
"prompt_tokens": 19,
"completion_tokens": 12,
"total_tokens": 31
},
"cumulative_logprob": [-2.045250159438183],
"finish_reason": "stop"
}
```

**When using vllm's OpenAI-compatible API service, it response complete results.**
```
python3.9", "-m", "vllm.entrypoints.openai.api_server", "--model", "/Qwen/Qwen1.5-72B-Chat", "--port", "21002", "--gpu-memory-utilization", "0.98", "--tensor-parallel-size", "4

{
"id": "cmpl-a10de9e3ba27482296eecc572b5623d6",
"object": "chat.completion",
"created": 443291,
"model": "/Qwen/Qwen1.5-72B-Chat",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I assist you today?\n"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 9,
"total_tokens": 21,
"completion_tokens": 12
}
}

```

**When using model_work, it response also complete results.**

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the API Server's /v1/chat/completions path and the vllm_worker /worker_generate_stream endpoint, then compare their handling of streamed results with vllm.entrypoints.openai.api_server. Reproduce using the Qwen1.5-72B-Chat command and the supplied curl requests. Done means the API returns the complete assistant response rather than the final truncated stream chunk, with ordering and finish_reason preserved.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
api, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.