abetlen / abetlen/llama-cpp-python

Docker GPU installation works in interactive mode but not in the dockerfile

Aperta
#742 2 commenti 1 reazione 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
10.6k
Fork
1.4k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

# Expected Behavior

I want llama-cpp-python to be able to load GGUF models with GPU inside docker. It works properly while installing llama-cpp-python on interactive mode but not inside the dockerfile. Since I work in a hospital my aim is to be able to do it offline (using the downloaded tar.gz file of llama-cpp-python).

# Environment and Context

I am working on a windows 11 machine and my docker container runs an ubuntu 20.04

# Failure Information (for bugs)

Basically while loading the model GPU is clearly not loaded (BLAS=0 and no messasge regarding layer offloading)

# Steps to Reproduce

I downloaded the tar.gz file of llama-cpp-python in a linux 20.04 via
```
pip download llama-cpp-python
```
Then I run the following inside the dockerfile:
```
RUN /bin/bash -c 'CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install /app/llama_cpp_folder/llama_cpp_python-0.2.6.tar.gz'
```
While trying to load the Llama model it does not load the GPU, it works, but without GPU.

However, if I go to interactive mode of the container, uninstall llama_cpp_python, and run the following command it works perfectly:
```
CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python
```

I also tried the following way to check whether the environment was not properly set and it does not load GPU either:
```
ENV CMAKE_ARGS="-DLLAMA_CUBLAS=on"
RUN pip install /app/llama_cpp_folder/llama_cpp_python-0.2.6.tar.gz
```

It seems to me that there is some problem regarding CUBLAS or CMAKE.

Thank you for your time,
Javi

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.