abetlen / abetlen/llama-cpp-python

When multiple requests are processed, the first request is interrupted

オープン
#867 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る
bug
主要言語
Python
スター
10.6k
フォーク
1.4k
PR マージ指標
PR 指標を取得中

説明

When multiple requests are processed, the first request is interrupted. How to solve this problem?

My run command is as follows:

python3 -m llama_cpp.server --model ./models/WizardLM-13B-V1.2/ggml-model-f16-Q5.gguf --n_gpu_layers 1 --n_ctx 8192

I tried to set up the following command:

python3 -m llama_cpp.server --model ./models/WizardLM-13B-V1.2/ggml-model-f16-Q5.gguf --n_gpu_layers 1 --n_ctx 8192 --interrupt_requests False

But --interrupt_requests False did not take effect.

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。