abetlen / abetlen/llama-cpp-python

GPU acceleration gives gibberish output and breaks string grammars

未關閉
#1,593 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
10.6k
分支
1.4k
平均合併
5 小時 23 分鐘
30 天內合併 PR
5

描述

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [x] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Model generates normal output when using GPU acceleration.

# Current Behavior

Model instead generates gibberish. Moreover, when using a grammar, the model encounters an assertion error.

For example, invoking the following code fragment:
```
from llama_cpp import Llama

llm = Llama(
model_path="./models/qwen2-7b-instruct-q5_k_m.gguf",
# n_gpu_layers=-1, # Uncomment to use GPU acceleration
seed=696969, # Uncomment to set a specific seed
n_ctx=32768, # Uncomment to increase the context window
)
output = llm(
"Q: Name the planets in the solar system? A: ", # Prompt
max_tokens=32, # Generate up to 32 tokens, set to None to generate up to the end of the context window
stop=["Q:", "\n"], # Stop generating just before the model would generate a new question
echo=True # Echo the prompt back in the output
) # Generate a completion, can also call create_completion
print(output)
```

With no GPU acceleration, the model produces the following output:
```
{'id': 'cmpl-8623a047-73a8-458b-9763-1a9f78f0fe04', 'object': 'text_completion', 'created': 1720787931, 'model': './models/qwen2-7b-instruct-q5_k_m.gguf', 'choices': [{'text': 'Q: Name the planets in the solar system? A: 1. Mercury 2. Venus 3. Earth 4. Mars 5. Jupiter 6. Saturn 7. Uranus 8. Neptune', 'index': 0, 'logprobs': None, 'finish_reason': 'length'}], 'usage': {'prompt_tokens': 13, 'completion_tokens': 32, 'total_tokens': 45}}
```

With GPU acceleration, the model produces the following output:
```
{'id': 'cmpl-b490a16d-d53b-4189-a5e6-40c4000425e9', 'object': 'text_completion', 'created': 1720787951, 'model': './models/qwen2-7b-instruct-q5_k_m.gguf', 'choices': [{'text': 'Q: Name the planets in the solar system? A: 1.GanG', 'index': 0, 'logprobs': None, 'finish_reason': 'stop'}], 'usage': {'prompt_tokens': 13, 'completion_tokens': 6, 'total_tokens': 19}}
```

Additionally, when using a string grammar, I encountered this assertion error:
```
GGML_ASSERT: C:\Users\minhk\AppData\Local\Temp\pip-install-h1m08tsi\llama-cpp-python_e9c0081b84634f459b14411750bdc6a0\vendor\llama.cpp\src\llama.cpp:17594: !grammar->stacks.empty()
```

# Environment and Context

My laptop has a NVIDIA GeForce GTX 4060 GPU and is running with CUDA 12.5.1.
OS: Windows 11 Home version 10.0.22621 build 22621.

Python version:
```
python --version
```
```
Python 3.12.1
```

Make version:
```
make --version
```
```
GNU Make 4.4.1
Built for Windows32
Copyright (C) 1988-2023 Free Software Foundation, Inc.
License GPLv3+: GNU GPL version 3 or later
This is free software: you are free to change and redistribute it.
There is NO WARRANTY, to the extent permitted by law.
```

GCC version:
```
g++ --version
```
```
g++.exe (MinGW-W64 x86_64-ucrt-posix-seh, built by Brecht Sanders) 12.3.0
Copyright (C) 2022 Free Software Foundation, Inc.
This is free software; see the source for copying conditions. There is NO
warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.
```

# Failure Information (for bugs)

Please help provide information about the failure if this is a bug. If it is not a bug, please remove the rest of this template.

# Steps to Reproduce

1. Install `llama-cpp-python` with CUDA.
2. Use [Qwen2-7B-Instruct](https://huggingface.co/Qwen/Qwen2-7B-Instruct) model

# Failure Logs

There are no failure logs - the program just returns nonsensical output.

貢獻指南

開啟貢獻指南

研究方向

問題出在 GPU 加速上,會導致輸出變成亂碼並破壞 grammar。先檢查 llama-cpp-python 的 CUDA 建置,具體查看 GPU 層卸載。查看檔案 vendor/llama.cpp/src/llama.cpp 約第 17594 行的 grammar stack 斷言。使用提供的 Qwen2 模型重現,並比較 CPU 與 GPU 推論路徑。檢查是否存在與 Windows CUDA 12.5 及特定 GPU 架構相關的已知問題。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
ai, backend, tooling
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
描述清楚
新手友好度
30/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。