abetlen / abetlen/llama-cpp-python

Multiple GPU incredibly slow inference

Đang mở
#1,026 6 bình luận 0 reaction 0 người được giao Xem trên GitHub
bug
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [ ] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Same or comparable inference speed on a single A100 vs 2 A100 setup.

# Current Behavior

GPU inference stats when all two GPUs are available to the inference process (30-60x) slower when compared to a single GPU run:

![image](https://github.com/abetlen/llama-cpp-python/assets/48142538/677728e1-1d34-4dd2-a299-84f0a1ddc49a)

The best solution i found is to manually hide the second GPU using CUDA_VISIBLE_DEVICES="0".

The model is initialized with main_gpu=0, tensor_split=None. In addition, when all 2 GPUs are visible, tensor_split option doesnt work as expected, since nvidia-smi shows, that both GPUs are used.

# Environment and Context

2x A100 GPU server, cuda 12.1, evaluated llama-cpp-python versions: 2.11, 2.13, 2.19 with cuBLAS backend

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.