abetlen / abetlen/llama-cpp-python
When multiple requests are processed, the first request is interrupted
未關閉
bug
- 主要語言
- Python
- 星號
- 10.6k
- 分支
- 1.4k
- PR 合併指標
- PR 指標待擷取
描述
When multiple requests are processed, the first request is interrupted. How to solve this problem?
My run command is as follows:
python3 -m llama_cpp.server --model ./models/WizardLM-13B-V1.2/ggml-model-f16-Q5.gguf --n_gpu_layers 1 --n_ctx 8192
I tried to set up the following command:
python3 -m llama_cpp.server --model ./models/WizardLM-13B-V1.2/ggml-model-f16-Q5.gguf --n_gpu_layers 1 --n_ctx 8192 --interrupt_requests False
But --interrupt_requests False did not take effect.
貢獻指南
評估
這個 Issue 還沒有評估資料。