abetlen / abetlen/llama-cpp-python
When multiple requests are processed, the first request is interrupted
未关闭
bug
- 主要语言
- Python
- 星标
- 10.6k
- 派生
- 1.4k
- PR 合并指标
- PR 指标待抓取
描述
When multiple requests are processed, the first request is interrupted. How to solve this problem?
My run command is as follows:
python3 -m llama_cpp.server --model ./models/WizardLM-13B-V1.2/ggml-model-f16-Q5.gguf --n_gpu_layers 1 --n_ctx 8192
I tried to set up the following command:
python3 -m llama_cpp.server --model ./models/WizardLM-13B-V1.2/ggml-model-f16-Q5.gguf --n_gpu_layers 1 --n_ctx 8192 --interrupt_requests False
But --interrupt_requests False did not take effect.
贡献指南
评估
这个 Issue 还没有评估数据。