abetlen / abetlen/llama-cpp-python

CPU Usage is very low when using it with BLAS

オープン
#811 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
bug
主要言語
Python
スター
10.6k
フォーク
1.4k
PR マージ指標
PR 指標を取得中

説明

Hello,

Thank you for the great work you guys are doing. 🙏

This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏

I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.

can you please guide me how i increase the speed of the inference. Any help would be appreciated.

Thank you 🙏

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。