abetlen / abetlen/llama-cpp-python

CUDA not supported. `ValueError: Attempt to split tensors that exceed maximum supported devices. Current LLAMA_MAX_DEVICES=1`

Abierto
#1,692 1 comentario 1 reacción 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
10.6k
Forks
1.4k
Métricas de merge de PR
Métricas de PR pendientes

Descripción

This was a problem that I think was prematurely `closed`:

https://github.com/abetlen/llama-cpp-python/issues/1166

My current efforts are to get a llama 3.1 70B gguf running on 2 3090s, and no matter my installation method, I'm getting the same error. Moreover, it appears `llama_cpp.llama_supports_gpu_offload()` always reports `False` even though it can use a single GPU.

Error:

```sh
# it never even hears the ENV VAR (!?), still reports as 1 device
$ LLAMA_MAX_DEVICES=2 my_thing
ValueError: Attempt to split tensors that exceed maximum supported devices. Current LLAMA_MAX_DEVICES=1
```

Installation Methods:

```sh
# 1
pip install llama-cpp-python \
--extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu122

# 2
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python

# 4
pip install llama-cpp-python --upgrade --force-reinstall --no-cache-dir

# 5
pip install llama-cpp-python --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu122 --upgrade --force-reinstall --no-cache-dir

# 6 downgrade was reported in this Issue to work, but does not
pip install llama-cpp-python==0.2.77 --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu122 --upgrade --force-reinstall

# 7
pip install llama-cpp-python==0.2.76 --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu122 --upgrade --force-reinstall

# 8
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python==0.2.77 --upgrade --force-reinstall --no-cache-dir --verbose

# 9 build fails
git checkout "v0.2.77"
CMAKE_ARGS="-DGGML_CUDA=on" pip install -e . --upgrade --force-reinstall --no-cache-dir --verbose

CMake Error at CMakeLists.txt:25 (add_subdirectory):
The source directory
llama-cpp-python/vendor/llama.cpp
does not contain a CMakeLists.txt file.

# 10 downgrade build fails the same
git clone ...
CMAKE_ARGS="-DGGML_CUDA=on" pip install -e ../lib/llama-cpp-python/ --verbose

# 11 maybe we copy that CMakeLists in? nope.
$ cp CMakeLists.txt vendor/llama.cpp/
$ CMAKE_ARGS="-DGGML_CUDA=on" pip install -e . --upgrade --force-reinstall --no-cache-dir --verbose

CMake Error at vendor/llama.cpp/CMakeLists.txt:25 (add_subdirectory):
add_subdirectory given source "vendor/llama.cpp" which is not an existing
directory.
```

After each attempt (besides builds, which all fail):

```sh
$ python -c "import llama_cpp; print(llama_cpp.llama_max_devices())"
1

$ python -c "import llama_cpp; print(llama_cpp.llama_supports_gpu_offload())"
False

$ python3 -c "import torch; print(torch.cuda.device_count())"
2
```

_Originally posted by @freckletonj in https://github.com/abetlen/llama-cpp-python/issues/1166#issuecomment-2294990187_

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.