abetlen / abetlen/llama-cpp-python
CPU Usage is very low when using it with BLAS
Offen
bug
- Vorherrschende Sprache
- Python
- Sterne
- 10.6k
- Forks
- 1.4k
- PR-Merge-Kennzahlen
- PR-Kennzahlen ausstehend
Beschreibung
Hello,
Thank you for the great work you guys are doing. 🙏
This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏
I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.
can you please guide me how i increase the speed of the inference. Any help would be appreciated.
Thank you 🙏
Beitragsleitfaden
Bewertung
Dieses Issue wurde noch nicht bewertet.