abetlen / abetlen/llama-cpp-python

CPU Usage is very low when using it with BLAS

Open
#811 0 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

Hello,

Thank you for the great work you guys are doing. 🙏

This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏

I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.

can you please guide me how i increase the speed of the inference. Any help would be appreciated.

Thank you 🙏

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.