abetlen / abetlen/llama-cpp-python

Docker GPU installation works in interactive mode but not in the dockerfile

Open
#742 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

# Expected Behavior

I want llama-cpp-python to be able to load GGUF models with GPU inside docker. It works properly while installing llama-cpp-python on interactive mode but not inside the dockerfile. Since I work in a hospital my aim is to be able to do it offline (using the downloaded tar.gz file of llama-cpp-python).

# Environment and Context

I am working on a windows 11 machine and my docker container runs an ubuntu 20.04

# Failure Information (for bugs)

Basically while loading the model GPU is clearly not loaded (BLAS=0 and no messasge regarding layer offloading)

# Steps to Reproduce

I downloaded the tar.gz file of llama-cpp-python in a linux 20.04 via
```
pip download llama-cpp-python
```
Then I run the following inside the dockerfile:
```
RUN /bin/bash -c 'CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install /app/llama_cpp_folder/llama_cpp_python-0.2.6.tar.gz'
```
While trying to load the Llama model it does not load the GPU, it works, but without GPU.

However, if I go to interactive mode of the container, uninstall llama_cpp_python, and run the following command it works perfectly:
```
CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python
```

I also tried the following way to check whether the environment was not properly set and it does not load GPU either:
```
ENV CMAKE_ARGS="-DLLAMA_CUBLAS=on"
RUN pip install /app/llama_cpp_folder/llama_cpp_python-0.2.6.tar.gz
```

It seems to me that there is some problem regarding CUBLAS or CMAKE.

Thank you for your time,
Javi

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.