abetlen / abetlen/llama-cpp-python
Docker GPU installation works in interactive mode but not in the dockerfile
- Ngôn ngữ chính
- Python
- Star
- 10.6k
- Fork
- 1.4k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
# Expected Behavior
I want llama-cpp-python to be able to load GGUF models with GPU inside docker. It works properly while installing llama-cpp-python on interactive mode but not inside the dockerfile. Since I work in a hospital my aim is to be able to do it offline (using the downloaded tar.gz file of llama-cpp-python).
# Environment and Context
I am working on a windows 11 machine and my docker container runs an ubuntu 20.04
# Failure Information (for bugs)
Basically while loading the model GPU is clearly not loaded (BLAS=0 and no messasge regarding layer offloading)
# Steps to Reproduce
I downloaded the tar.gz file of llama-cpp-python in a linux 20.04 via
```
pip download llama-cpp-python
```
Then I run the following inside the dockerfile:
```
RUN /bin/bash -c 'CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install /app/llama_cpp_folder/llama_cpp_python-0.2.6.tar.gz'
```
While trying to load the Llama model it does not load the GPU, it works, but without GPU.
However, if I go to interactive mode of the container, uninstall llama_cpp_python, and run the following command it works perfectly:
```
CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python
```
I also tried the following way to check whether the environment was not properly set and it does not load GPU either:
```
ENV CMAKE_ARGS="-DLLAMA_CUBLAS=on"
RUN pip install /app/llama_cpp_folder/llama_cpp_python-0.2.6.tar.gz
```
It seems to me that there is some problem regarding CUBLAS or CMAKE.
Thank you for your time,
Javi
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.