abetlen / abetlen/llama-cpp-python
CPU Usage is very low when using it with BLAS
Ouverte
bug
- Langage dominant
- Python
- Étoiles
- 10.6k
- Forks
- 1.4k
- Métriques de merge des PR
- Métriques de PR en attente
Description
Hello,
Thank you for the great work you guys are doing. 🙏
This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏
I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.
can you please guide me how i increase the speed of the inference. Any help would be appreciated.
Thank you 🙏
Guide de contribution
Ouvrir le guide de contribution
Évaluation
Cette issue n'a pas encore été évaluée.