fastchat.serve.model_worker --device cpu only uses one CPU Thread for token generation.
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
i launch a worker with python3 -m fastchat.serve.model_worker --model-path /home/llamaweights/vicuna-13b --device cpu and then the webGUI, which works fine so far. When i do a request, after an initial loading time one core goes to 100% while the others idle. If i make a second request in another tap another core goes to 100% while the other 14 idle... Token generation is very slow, but does not get any slower for additional requests. Can i somehow use all 16 threads or at least all 8 cores for a single request to speed up token generation?
Kind regards
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the behavior with python3 -m fastchat.serve.model_worker --model-path /home/llamaweights/vicuna-13b --device cpu and inspect the CPU execution path used by fastchat.serve.model_worker. Done means determining whether one request can use the available CPU threads and documenting or implementing a supported way to achieve that, if applicable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100