abetlen / abetlen/llama-cpp-python

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1  behavior is strange.

Đang mở
#1,720 3 bình luận 2 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [ ] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [ ] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [ ] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [ ] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Prioritize use of VRAM, and start using shared memory when memory is exceeded
and
Fast inference

# Current Behavior

export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
When you use this option, RAM will be used first instead of VRAM.
Also, the specified GPU will not be used first.
`llama_print_timings: total time = 56361.73 ms / 45 tokens`

Hiding the option makes it super fast
`llama_print_timings: total time = 40.95 ms / 143 tokens`

# Environment and Context

```
Windows11 WSL2 Ubuntu 22.04.4 LTS
CUDA12.1

Python 3.10.11
GNU Make 4.3 x86_64-pc-linux-gnu
g++ (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
```

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.