abetlen / abetlen/llama-cpp-python

Dynamically intterupt token generation

Đang mở
#599 0 bình luận 2 reaction 0 người được giao Xem trên GitHub
enhancement
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

**Is your feature request related to a problem? Please describe.**
During the generation of tokens I would like to stop when I encounter some condition that changes during the runtime, using `Stream=True`.
E.g. I would like to stop generation after 5 lines of generation.

**Describe the solution you'd like**
I would like a method on llm called `stop()`, or `interrupt()`, that forces the model to stop after the next token is generated, similar to CTRL+C in the regular llama.cpp

**Describe alternatives you've considered**
I have considered adding a newline as stop token, but I think this is not performant. Another way I can think of is changing the `stop` list after passing it the generation method, but that feels hacky.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.