lm-sys / lm-sys/FastChat

AYA-101 killing SGLang

Open
#3,049 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

I tried [Aya-101](https://huggingface.co/CohereForAI/aya-101), the multilingual model, with sglang worker, and I get this. Maybe it happens to other models as well?

```
2024-02-15 16:04:21 | INFO | stdout | router init state: Traceback (most recent call last):
2024-02-15 16:04:21 | INFO | stdout | File "/p/haicluster/llama/FastChat/sc_venv_2024/venv/lib/python3.11/site-packages/sglang/srt/managers/router/manager.py", line 68, in start_router_process
2024-02-15 16:04:21 | INFO | stdout | model_client = ModelRpcClient(server_args, port_args)
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/p/haicluster/llama/FastChat/sc_venv_2024/venv/lib/python3.11/site-packages/sglang/srt/managers/router/model_rpc.py", line 564, in __init__
2024-02-15 16:04:21 | INFO | stdout | self.model_servers = [x[0] for x in rets]
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/p/haicluster/llama/FastChat/sc_venv_2024/venv/lib/python3.11/site-packages/sglang/srt/managers/router/model_rpc.py", line 564, in
2024-02-15 16:04:21 | INFO | stdout | self.model_servers = [x[0] for x in rets]
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/concurrent/futures/_base.py", line 619, in result_iterator
2024-02-15 16:04:21 | INFO | stdout | yield _result_or_cancel(fs.pop())
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/concurrent/futures/_base.py", line 317, in _result_or_cancel
2024-02-15 16:04:21 | INFO | stdout | return fut.result(timeout)
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/concurrent/futures/_base.py", line 456, in result
2024-02-15 16:04:21 | INFO | stdout | return self.__get_result()
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/concurrent/futures/_base.py", line 401, in __get_result
2024-02-15 16:04:21 | INFO | stdout | raise self._exception
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/concurrent/futures/thread.py", line 58, in run
2024-02-15 16:04:21 | INFO | stdout | result = self.fn(*self.args, **self.kwargs)
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/p/haicluster/llama/FastChat/sc_venv_2024/venv/lib/python3.11/site-packages/sglang/srt/managers/router/model_rpc.py", line 597, in start_model_process
2024-02-15 16:04:21 | INFO | stdout | proc.start()
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/process.py", line 121, in start
2024-02-15 16:04:21 | INFO | stdout | self._popen = self._Popen(self)
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/context.py", line 224, in _Popen
2024-02-15 16:04:21 | INFO | stdout | return _default_context.get_context().Process._Popen(process_obj)
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/context.py", line 288, in _Popen
2024-02-15 16:04:21 | INFO | stdout | return Popen(process_obj)
2024-02-15 16:04:21 | INFO | stdout | ^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/popen_spawn_posix.py", line 32, in __init__
2024-02-15 16:04:21 | INFO | stdout | super().__init__(process_obj)
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/popen_fork.py", line 19, in __init__
2024-02-15 16:04:21 | INFO | stdout | self._launch(process_obj)
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/popen_spawn_posix.py", line 47, in _launch
2024-02-15 16:04:21 | INFO | stdout | reduction.dump(process_obj, fp)
2024-02-15 16:04:21 | INFO | stdout | File "/easybuild/2024/software/Python/3.11.3-GCCcore-12.3.0/lib/python3.11/multiprocessing/reduction.py", line 60, in dump
2024-02-15 16:04:21 | INFO | stdout | ForkingPickler(file, protocol).dump(obj)
2024-02-15 16:04:21 | INFO | stdout | AttributeError: Can't pickle local object 'start_model_process.._init_service'
2024-02-15 16:04:21 | INFO | stdout |
2024-02-15 16:04:21 | INFO | stdout | detoken init state: init ok
2024-02-15 16:04:22 | ERROR | stderr | Traceback (most recent call last):
2024-02-15 16:04:22 | ERROR | stderr | File "/p/haicluster/llama/FastChat/fastchat/serve/sglang_worker.py", line 290, in
2024-02-15 16:04:22 | ERROR | stderr | runtime = sgl.Runtime(
2024-02-15 16:04:22 | ERROR | stderr | ^^^^^^^^^^^^
2024-02-15 16:04:22 | ERROR | stderr | File "/p/haicluster/llama/FastChat/sc_venv_2024/venv/lib/python3.11/site-packages/sglang/api.py", line 39, in Runtime
2024-02-15 16:04:22 | ERROR | stderr | return Runtime(*args, **kwargs)
2024-02-15 16:04:22 | ERROR | stderr | ^^^^^^^^^^^^^^^^^^^^^^^^
2024-02-15 16:04:22 | ERROR | stderr | File "/p/haicluster/llama/FastChat/sc_venv_2024/venv/lib/python3.11/site-packages/sglang/srt/server.py", line 482, in __init__
2024-02-15 16:04:22 | ERROR | stderr | raise RuntimeError("Launch failed. Please see the error messages above.")
2024-02-15 16:04:22 | ERROR | stderr | RuntimeError: Launch failed. Please see the error messages above.
srun: error: haicluster1: task 0: Exited with exit code 1
```

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at fastchat/serve/sglang_worker.py around line 290, then inspect the sglang router/model_rpc.py traceback around start_model_process. Reproduce the Aya-101 launch and investigate the reported multiprocessing pickle failure. Done means the sglang worker launches successfully without the Launch failed error.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.