abetlen / abetlen/llama-cpp-python

Feature request: ability to tokenize a list of strings _or_ keep the tokenizer warm

Đang mở
#1,763 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
enhancement
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

*Situation:* let's say you have a list of sentences you want to tokenize.

*Current workaround:*
```python
from llama_cpp import Llama

embedder = Llama.from_pretrained(repo_id="lm-kit/bge-m3-gguf", filename="*F16.gguf", embedding=True)
sentences = ["Hello world"] * 1000
sentences_tokens = [embedder.tokenize(sentence.encode()) for sentence in sentences]
# ↑ Each tokenize call has an overhead of 200-400ms
```

*Problem:* each call to `tokenize` appears to have an overhead of about 200-400ms. Which means tokenizing 1000 sentences will take 200-400 seconds 💥. Even tokenizing 10 sentences can take 4 seconds!

*Feature request*: either add the ability to tokenize a list of strings efficiently with `tokenize`, or add the ability to keep the tokenizer warm so that subsequent calls are not as slow as the first call.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.