abetlen / abetlen/llama-cpp-python

Include usage key in create_completion when streaming

オープン
#1,498 コメント 2 件 リアクション 4 件 担当者 0 名 GitHub で見る
enhancement
主要言語
Python
スター
10.6k
フォーク
1.4k
PR マージ指標
PR 指標を取得中

説明

**Is your feature request related to a problem? Please describe.**
Since `create_completion` may yield text chunks comprised of multiple tokens per yield (e.g. in the case of multi-byte Unicode characters), counting the number of yields may not equal the number of tokens actually generated by a model. To accurately get the usage statistics of a streamed completion, one has to run the final text through the tokenizer again, despite `create_completion` already tracking the number of tokens generated by the model.

**Describe the solution you'd like**
When `stream=True` in `create_completion`, the final chunk yielded should include the usage statistics in the `'usage'` key.

**Describe alternatives you've considered**
- Saving full generated text and running it through the tokenizer again (seems wasteful)
- Counting the number of yields and hoping we don't have any multi-byte characters (hacky and fragile)

**Additional context**
The OpenAI API has recently added similar support in their streaming API with the `stream_options` key: https://platform.openai.com/docs/api-reference/chat/create#chat-create-stream_options

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。