abetlen / abetlen/llama-cpp-python

CPU Usage is very low when using it with BLAS

Aberta
#811 0 comentários 0 reações 0 responsáveis Ver no GitHub
bug
Linguagem predominante
Python
Estrelas
10.6k
Forks
1.4k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

Hello,

Thank you for the great work you guys are doing. 🙏

This is my first time posting the issue here so if i make any mistake apologies in advance. 🙏

I am using llama-cpp-python with BLAS on CPU as i don't have access to GPU. The CPU usage is only 50% and generation time is very slow for the longer prompt mine is around 200 tokens long prompt.

can you please guide me how i increase the speed of the inference. Any help would be appreciated.

Thank you 🙏

Guia de contribuição

Abrir o guia de contribuição

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.