abetlen / abetlen/llama-cpp-python
CPU Usage is very low when using it with BLAS
Aberta
bug
- Linguagem predominante
- Python
- Estrelas
- 10.6k
- Forks
- 1.4k
- Métricas de merge de PRs
- Métricas de PR pendentes
Descrição
Hello,
Thank you for the great work you guys are doing. 🙏
This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏
I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.
can you please guide me how i increase the speed of the inference. Any help would be appreciated.
Thank you 🙏
Guia de contribuição
Avaliação
Esta issue ainda não foi avaliada.