abetlen / abetlen/llama-cpp-python
CPU Usage is very low when using it with BLAS
Abierto
bug
- Lenguaje dominante
- Python
- Estrellas
- 10.6k
- Forks
- 1.4k
- Métricas de merge de PR
- Métricas de PR pendientes
Descripción
Hello,
Thank you for the great work you guys are doing. 🙏
This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏
I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.
can you please guide me how i increase the speed of the inference. Any help would be appreciated.
Thank you 🙏
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.