abetlen / abetlen/llama-cpp-python

Feature request: ability to tokenize a list of strings _or_ keep the tokenizer warm

Aperta
#1,763 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
enhancement
Lingua principale
Python
Stelle
10.6k
Fork
1.4k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

*Situation:* let's say you have a list of sentences you want to tokenize.

*Current workaround:*
```python
from llama_cpp import Llama

embedder = Llama.from_pretrained(repo_id="lm-kit/bge-m3-gguf", filename="*F16.gguf", embedding=True)
sentences = ["Hello world"] * 1000
sentences_tokens = [embedder.tokenize(sentence.encode()) for sentence in sentences]
# ↑ Each tokenize call has an overhead of 200-400ms
```

*Problem:* each call to `tokenize` appears to have an overhead of about 200-400ms. Which means tokenizing 1000 sentences will take 200-400 seconds 💥. Even tokenizing 10 sentences can take 4 seconds!

*Feature request*: either add the ability to tokenize a list of strings efficiently with `tokenize`, or add the ability to keep the tokenizer warm so that subsequent calls are not as slow as the first call.

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.