abetlen / abetlen/llama-cpp-python
Feature request: ability to tokenize a list of strings _or_ keep the tokenizer warm
- Dominant language
- Python
- Stars
- 10.6k
- Forks
- 1.4k
- PR merge metrics
- PR metrics pending
Description
*Situation:* let's say you have a list of sentences you want to tokenize.
*Current workaround:*
```python
from llama_cpp import Llama
embedder = Llama.from_pretrained(repo_id="lm-kit/bge-m3-gguf", filename="*F16.gguf", embedding=True)
sentences = ["Hello world"] * 1000
sentences_tokens = [embedder.tokenize(sentence.encode()) for sentence in sentences]
# ↑ Each tokenize call has an overhead of 200-400ms
```
*Problem:* each call to `tokenize` appears to have an overhead of about 200-400ms. Which means tokenizing 1000 sentences will take 200-400 seconds 💥. Even tokenizing 10 sentences can take 4 seconds!
*Feature request*: either add the ability to tokenize a list of strings efficiently with `tokenize`, or add the ability to keep the tokenizer warm so that subsequent calls are not as slow as the first call.
Contributor guide
Assessment
This issue has not been assessed yet.