abetlen / abetlen/llama-cpp-python
Multiple GPU incredibly slow inference
- Dominant language
- Python
- Stars
- 10.6k
- Forks
- 1.4k
- PR merge metrics
- PR metrics pending
Description
# Prerequisites
Please answer the following questions for yourself before submitting an issue.
- [ ] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.
# Expected Behavior
Same or comparable inference speed on a single A100 vs 2 A100 setup.
# Current Behavior
GPU inference stats when all two GPUs are available to the inference process (30-60x) slower when compared to a single GPU run:

The best solution i found is to manually hide the second GPU using CUDA_VISIBLE_DEVICES="0".
The model is initialized with main_gpu=0, tensor_split=None. In addition, when all 2 GPUs are visible, tensor_split option doesnt work as expected, since nvidia-smi shows, that both GPUs are used.
# Environment and Context
2x A100 GPU server, cuda 12.1, evaluated llama-cpp-python versions: 2.11, 2.13, 2.19 with cuBLAS backend
Contributor guide
Assessment
This issue has not been assessed yet.