abetlen / abetlen/llama-cpp-python

CPU Usage is very low when using it with BLAS

未关闭
#811 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
bug
主要语言
Python
星标
10.6k
派生
1.4k
PR 合并指标
PR 指标待抓取

描述

Hello,

Thank you for the great work you guys are doing. 🙏

This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏

I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.

can you please guide me how i increase the speed of the inference. Any help would be appreciated.

Thank you 🙏

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。