abetlen / abetlen/llama-cpp-python

CPU Usage is very low when using it with BLAS

Aperta
#811 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
bug
Lingua principale
Python
Stelle
10.6k
Fork
1.4k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Hello,

Thank you for the great work you guys are doing. 🙏

This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏

I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.

can you please guide me how i increase the speed of the inference. Any help would be appreciated.

Thank you 🙏

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.