abetlen / abetlen/llama-cpp-python

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1  behavior is strange.

オープン
#1,720 コメント 3 件 リアクション 2 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
10.6k
フォーク
1.4k
PR マージ指標
PR 指標を取得中

説明

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [ ] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [ ] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [ ] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [ ] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Prioritize use of VRAM, and start using shared memory when memory is exceeded
and
Fast inference

# Current Behavior

export GGML_CUDA_ENABLE_UNIFIED_MEMORY=1
When you use this option, RAM will be used first instead of VRAM.
Also, the specified GPU will not be used first.
`llama_print_timings: total time = 56361.73 ms / 45 tokens`

Hiding the option makes it super fast
`llama_print_timings: total time = 40.95 ms / 143 tokens`

# Environment and Context

```
Windows11 WSL2 Ubuntu 22.04.4 LTS
CUDA12.1

Python 3.10.11
GNU Make 4.3 x86_64-pc-linux-gnu
g++ (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
```

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。