abetlen / abetlen/llama-cpp-python

Enable / Add llama.cpp.python server --parallel | -np N | --n_parallel N

未關閉
#1,329 3 則留言 8 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
10.6k
分支
1.4k
PR 合併指標
PR 指標待擷取

描述

**Is your feature request related to a problem? Please describe.**
A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

When executing chat completions I need to wait for a prompt to complete before executing a new one. I'd like to be able to execute multiple prompts at the same time. Right now my GPU utilization on a g4dn.2xlarge instance is max 65-80% (model loaded in GPU mem) I've tinkered around with n_batch, ctx and a few other parameters.

**Describe the solution you'd like**
A clear and concise description of what you want to happen.
It seems like llama.cpp has this feature already -np N, --parallel N: Set the number of slots for process requests. Default: 1
https://github.com/ggerganov/llama.cpp/tree/master/examples/server

**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.

**Additional context**
Add any other context or screenshots about the feature request here.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。