abetlen / abetlen/llama-cpp-python
CPU Usage is very low when using it with BLAS
オープン
bug
- 主要言語
- Python
- スター
- 10.6k
- フォーク
- 1.4k
- PR マージ指標
- PR 指標を取得中
説明
Hello,
Thank you for the great work you guys are doing. 🙏
This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏
I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.
can you please guide me how i increase the speed of the inference. Any help would be appreciated.
Thank you 🙏
コントリビューションガイド
評価
この issue はまだ評価されていません。