abetlen / abetlen/llama-cpp-python

Include usage key in create_completion when streaming

Đang mở
#1,498 2 bình luận 4 reaction 0 người được giao Xem trên GitHub
enhancement
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

**Is your feature request related to a problem? Please describe.**
Since `create_completion` may yield text chunks comprised of multiple tokens per yield (e.g. in the case of multi-byte Unicode characters), counting the number of yields may not equal the number of tokens actually generated by a model. To accurately get the usage statistics of a streamed completion, one has to run the final text through the tokenizer again, despite `create_completion` already tracking the number of tokens generated by the model.

**Describe the solution you'd like**
When `stream=True` in `create_completion`, the final chunk yielded should include the usage statistics in the `'usage'` key.

**Describe alternatives you've considered**
- Saving full generated text and running it through the tokenizer again (seems wasteful)
- Counting the number of yields and hoping we don't have any multi-byte characters (hacky and fragile)

**Additional context**
The OpenAI API has recently added similar support in their streaming API with the `stream_options` key: https://platform.openai.com/docs/api-reference/chat/create#chat-create-stream_options

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.