lm-sys / lm-sys/FastChat

fastchat.serve.vllm_worker方法部署Yi-34B-Chat-4bits,实体抽取类任务 请求服务失败,内部错误

Open
#2,831 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

-----------------部署Yi-34B-Chat-4bits模型 ----------------

首先启动 controller :

nohup python3 -m fastchat.serve.controller --host 0.0.0.0 --port 21001 > ./log/controller.log 2>&1 &

启动 openapi的 兼容服务 地址 8000

nohup python3 -m fastchat.serve.openai_api_server --controller-address http://127.0.0.1:21001
--host 0.0.0.0 --port 8000 > ./log/api_server.log 2>&1 &

启动 web ui

nohup python -m fastchat.serve.gradio_web_server --controller-address http://127.0.0.1:21001
--host 0.0.0.0 --port 8000 > ./log/web_server.log 2>&1 &

启动模型: 说明,必须是本地ip --load-8bit 本身已经是int4了

nohup python3 -m fastchat.serve.vllm_worker --quantization awq --model-names yi-34b
--model-path /home/bigdata/model/01ai/Yi-34B-Chat-4bits --controller-address http://127.0.0.1:21001
--worker-address http://127.0.0.1:8080 --host 0.0.0.0 --port 8080
--num-gpus 2 > ./log/model_worker.log 2>&1 &

curl http://0.0.0.0:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model": "yi-34b","messages": [{"role": "user", "content": "北京景点,使用中文回答"}],"temperature": 0.7}'
以上case能够正常请求并得到结果

但是从通话文本中抽取实体任务,会发生内部错误
curl http://0.0.0.0:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model": "yi-34b","messages": [{"role": "user", "content": "通话文本: 您好,请问是韩先生家对吗?喂。对。我这边呢,是河北御芝林务中心的安排给您发货,可以吗?安排发货过去。啊,您这边订购的是变通胶囊,60装两瓶,一共是153块钱,对吗?对。 对。对。啊那同意就发货了啊。地址是黑龙江省牡丹江市西安区牡丹城32号楼。呃,发吧。2单元602对吗?605。啊,605对吗?嗯,改一下啊,好嘞,605对吗啊?牡丹城32号楼2单元605对 吗?605。605。嗯。嗯嗯。对对对。姓名姓韩韩先生,对吗?对对对。啊,那就发货了,三天左右到您这儿,您准备好153块钱给快递师傅就可以收货了啊。哎,行哎ok。呃,走的是圆通快 递啊。御芝林温馨提示您。保健食品不能代替药物,您购买的变通胶囊。保健功能的通便的。如果客服专员小刘夸大产品功能,或者家人对您本次购买不理解,你告诉我吗?我可以帮您终止订单,您这边没有疑问的话那发货了啊。啊丫头。嗯。啊,我妈吃了十天了。哦。是不是要把他一天都拉的便便三四渐变的,原来的便10年了。哦。嗯,之前是不是用了好好多,主要就不解决问题。嗯,什么送给他用了?多少树叶?常见了,这么大家都。哦,那行,嗯,那变动。不要用着用着挺快i。\n\n从上面的通话文本中提取顾客的身体病症,并以 JSON 格式提供,其中包含以 下键:custom_disease(你只需要回复病症)"}],"temperature": 0.7}'
错误栈如下:
2023-12-18 17:40:22 | INFO | stdout | INFO: 127.0.0.1:47468 - "POST /model_details HTTP/1.1" 200 OK
2023-12-18 17:40:22 | INFO | stdout | INFO: 127.0.0.1:47470 - "POST /count_token HTTP/1.1" 200 OK
INFO 12-18 17:40:22 async_llm_engine.py:379] Received request 52d69e4c79314bf9b0b67519968590f1: prompt: '<|im_start|>user\n通话文本: 您好,请问是韩先生家对吗?喂。对。我这边呢,是河北御芝林务中心的安排给您发货,可以吗?安排发货过去。啊,您这边订购的是变通胶囊,60装两瓶,一共是153块钱,对吗?对。 对。对。啊那同意就发货了啊。地址是黑龙江省牡丹江市西安区牡丹城32号楼。呃,发吧。2单元602对吗?605。啊,605对吗?嗯,改一下啊,好嘞,605对吗啊?牡丹城32号楼2单元605对 吗?605。605。嗯。嗯嗯。对对对。姓名姓韩韩先生,对吗?对对对。啊,那就发货了,三天左右到您这儿,您准备好153块钱给快递师傅就可以收货了啊。哎,行哎ok。呃,走的是圆通快 递啊。御芝林温馨提示您。保健食品不能代替药物,您购买的变通胶囊。保健功能的通便的。如果客服专员小刘夸大产品功能,或者家人对您本次购买不理解,你告诉我吗?我可以帮您终止订单,您这边没有疑问的话那发货了啊。啊丫头。嗯。啊,我妈吃了十天了。哦。是不是要把他一天都拉的便便三四渐变的,原来的便10年了。哦。嗯,之前是不是用了好好多,主要就不解决问题。嗯,什么送给他用了?多少树叶?常见了,这么大家都。哦,那行,嗯,那变动。不要用着用着挺快i。\n\n从上面的通话文本中提取顾客的身体病症,并以 JSON 格式提供,其中包含以下键:custom_disease(你只需要回复病症)<|im_end|>\n<|im_start|>assistant\n', sampling params: SamplingParams(n=1, best_of=1, presence_penalty=0.0, frequency_penalty=0.0, repetition_penalty=1.0, temperature=0.7, top_p=1.0, top_k=-1, min_p=0.0, use_beam_search=False, length_penalty=1.0, early_stopping=False, stop=['<|im_end|>', '<|im_start|>', '<|im_sep|>', '<|endoftext|>', ''], stop_token_ids=[2, 6, 7, 8, 64001], include_stop_str_in_output=False, ignore_eos=False, max_tokens=3677, logprobs=None, prompt_logprobs=None, skip_special_tokens=True, spaces_between_special_tokens=True), prompt token ids: None.
2023-12-18 17:40:24 | INFO | model_worker | Send heart beat. Models: ['yi-34b']. Semaphore: Semaphore(value=962, locked=False). call_ct: 79. worker_id: 66b79766.
2023-12-18 17:40:25 | ERROR | asyncio | Exception in callback functools.partial(<function _raise_exception_on_finish at 0x7f17fe4c0700>, request_tracker=<vllm.engine.async_llm_engine.RequestTracker object at 0x7f17e49ef220>)
handle: <Handle functools.partial(<function _raise_exception_on_finish at 0x7f17fe4c0700>, request_tracker=<vllm.engine.async_llm_engine.RequestTracker object at 0x7f17e49ef220>)>
Traceback (most recent call last):
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 28, in _raise_exception_on_finish
task.result()
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 359, in run_engine_loop
has_requests_in_progress = await self.engine_step()
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 338, in engine_step
request_outputs = await self.engine.step_async()
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 191, in step_async
output = await self._run_workers_async(
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 227, in _run_workers_async
assert output == other_output
AssertionError

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
File "uvloop/cbhandles.pyx", line 63, in uvloop.loop.Handle._run
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 37, in _raise_exception_on_finish
raise exc
File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 32, in _raise_exception_on_finish
raise AsyncEngineDeadError(
vllm.engine.async_llm_engine.AsyncEngineDeadError: Task finished unexpectedly. This should never happen! Please open an issue on Github. See stack trace above for the actual cause.
INFO 12-18 17:40:25 async_llm_engine.py:134] Aborted request 52d69e4c79314bf9b0b67519968590f1.
2023-12-18 17:40:25 | INFO | stdout | INFO: 127.0.0.1:47472 - "POST /worker_generate HTTP/1.1" 500 Internal Server Error
2023-12-18 17:40:25 | ERROR | stderr | ERROR: Exception in ASGI application
2023-12-18 17:40:25 | ERROR | stderr | Traceback (most recent call last):
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 28, in _raise_exception_on_finish
2023-12-18 17:40:25 | ERROR | stderr | task.result()
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 359, in run_engine_loop
2023-12-18 17:40:25 | ERROR | stderr | has_requests_in_progress = await self.engine_step()
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 338, in engine_step
2023-12-18 17:40:25 | ERROR | stderr | request_outputs = await self.engine.step_async()
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 191, in step_async
2023-12-18 17:40:25 | ERROR | stderr | output = await self._run_workers_async(
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 227, in _run_workers_async
2023-12-18 17:40:25 | ERROR | stderr | assert output == other_output
2023-12-18 17:40:25 | ERROR | stderr | AssertionError
2023-12-18 17:40:25 | ERROR | stderr |
2023-12-18 17:40:25 | ERROR | stderr | The above exception was the direct cause of the following exception:
2023-12-18 17:40:25 | ERROR | stderr |
2023-12-18 17:40:25 | ERROR | stderr | Traceback (most recent call last):
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/uvicorn/protocols/http/httptools_impl.py", line 426, in run_asgi
2023-12-18 17:40:25 | ERROR | stderr | result = await app( # type: ignore[func-returns-value]
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/uvicorn/middleware/proxy_headers.py", line 84, in call
2023-12-18 17:40:25 | ERROR | stderr | return await self.app(scope, receive, send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastapi/applications.py", line 1106, in call
2023-12-18 17:40:25 | ERROR | stderr | await super().call(scope, receive, send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/applications.py", line 122, in call
2023-12-18 17:40:25 | ERROR | stderr | await self.middleware_stack(scope, receive, send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/middleware/errors.py", line 184, in call
2023-12-18 17:40:25 | ERROR | stderr | raise exc
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/middleware/errors.py", line 162, in call
2023-12-18 17:40:25 | ERROR | stderr | await self.app(scope, receive, _send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/middleware/exceptions.py", line 79, in call
2023-12-18 17:40:25 | ERROR | stderr | raise exc
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/middleware/exceptions.py", line 68, in call
2023-12-18 17:40:25 | ERROR | stderr | await self.app(scope, receive, sender)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastapi/middleware/asyncexitstack.py", line 20, in call
2023-12-18 17:40:25 | ERROR | stderr | raise e
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastapi/middleware/asyncexitstack.py", line 17, in call
2023-12-18 17:40:25 | ERROR | stderr | await self.app(scope, receive, send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/routing.py", line 718, in call
2023-12-18 17:40:25 | ERROR | stderr | await route.handle(scope, receive, send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/routing.py", line 276, in handle
2023-12-18 17:40:25 | ERROR | stderr | await self.app(scope, receive, send)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/starlette/routing.py", line 66, in app
2023-12-18 17:40:25 | ERROR | stderr | response = await func(request)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastapi/routing.py", line 274, in app
2023-12-18 17:40:25 | ERROR | stderr | raw_response = await run_endpoint_function(
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastapi/routing.py", line 191, in run_endpoint_function
2023-12-18 17:40:25 | ERROR | stderr | return await dependant.call(**values)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastchat/serve/vllm_worker.py", line 196, in api_generate
2023-12-18 17:40:25 | ERROR | stderr | output = await worker.generate(params)
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastchat/serve/vllm_worker.py", line 154, in generate
2023-12-18 17:40:25 | ERROR | stderr | async for x in self.generate_stream(params):
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/fastchat/serve/vllm_worker.py", line 113, in generate_stream
2023-12-18 17:40:25 | ERROR | stderr | async for request_output in results_generator:
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 445, in generate
2023-12-18 17:40:25 | ERROR | stderr | raise e
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 439, in generate
2023-12-18 17:40:25 | ERROR | stderr | async for request_output in stream:
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 70, in anext
2023-12-18 17:40:25 | ERROR | stderr | raise result
2023-12-18 17:40:25 | ERROR | stderr | File "uvloop/cbhandles.pyx", line 63, in uvloop.loop.Handle._run
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 37, in _raise_exception_on_finish
2023-12-18 17:40:25 | ERROR | stderr | raise exc
2023-12-18 17:40:25 | ERROR | stderr | File "/home/bigdata/micromamba/envs/fastchat/lib/python3.10/site-packages/vllm/engine/async_llm_engine.py", line 32, in _raise_exception_on_finish
2023-12-18 17:40:25 | ERROR | stderr | raise AsyncEngineDeadError(
2023-12-18 17:40:25 | ERROR | stderr | vllm.engine.async_llm_engine.AsyncEngineDeadError: Task finished unexpectedly. This should never happen! Please open an issue on Github. See stack trace above for the actual cause.

请帮我确认下具体错误原因以及如何解决
thanks

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with fastchat/serve/vllm_worker.py and the vllm async_llm_engine.py stack trace, then reproduce the two curl requests using the Yi-34B-Chat-4bits setup and two GPUs. Compare the successful request with the entity-extraction request around the worker output assertion. Done means the failing request no longer returns HTTP 500 and the engine remains usable.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
api, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.