abetlen / abetlen/llama-cpp-python

Dynamically intterupt token generation

Open
#599 0 comments 2 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

**Is your feature request related to a problem? Please describe.**
During the generation of tokens I would like to stop when I encounter some condition that changes during the runtime, using `Stream=True`.
E.g. I would like to stop generation after 5 lines of generation.

**Describe the solution you'd like**
I would like a method on llm called `stop()`, or `interrupt()`, that forces the model to stop after the next token is generated, similar to CTRL+C in the regular llama.cpp

**Describe alternatives you've considered**
I have considered adding a newline as stop token, but I think this is not performant. Another way I can think of is changing the `stop` list after passing it the generation method, but that feels hacky.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.