abetlen / abetlen/llama-cpp-python

Running basic example from docs results in `TypeError: 'NoneType' object is not callable`

未关闭
#1,998 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
10.6k
派生
1.4k
PR 合并指标
PR 指标待抓取

描述

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [x] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

I am trying to run the most basic examples [from the docs][1].

I am trying to load a model that I've downloaded:

```python
from llama_cpp import Llama

if __name__ == "__main__":
llm = Llama(
model_path="models/qwen2-0_5b-instruct-q4_0.gguf",
verbose=False,
)
output = llm(
"Q: Name the planets in the solar system? A: ",
max_tokens=32,
stop=["Q:", "\n"],
echo=True,
)
print(output)
```

And I am also trying to load a model directly from Hugging Face hub:

```python
from llama_cpp import Llama

if __name__ == "__main__":
llm = Llama.from_pretrained(
repo_id="Qwen/Qwen2-0.5B-Instruct-GGUF",
filename="*q4_0.gguf",
verbose=False,
)
```

I expect these basic examples to run without error and produce some reasonable looking output.

[1]: https://llama-cpp-python.readthedocs.io/en/latest/#high-level-api

# Current Behavior

In both cases -- whether using a pre-downloaded model or pulling from Hugging Face Hub -- I am getting the following exception:

```python
Exception ignored in:
Traceback (most recent call last):
File ".../.venv/lib/python3.11/site-packages/llama_cpp/llama.py", line 2205, in __del__
File ".../.venv/lib/python3.11/site-packages/llama_cpp/llama.py", line 2202, in close
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 609, in close
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 601, in __exit__
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 586, in __exit__
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 360, in __exit__
File ".../.venv/lib/python3.11/site-packages/llama_cpp/_internals.py", line 75, in close
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 609, in close
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 601, in __exit__
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 586, in __exit__
File ".../.pyenv/versions/3.11.11/lib/python3.11/contextlib.py", line 469, in _exit_wrapper
File ".../.venv/lib/python3.11/site-packages/llama_cpp/_internals.py", line 69, in free_model
TypeError: 'NoneType' object is not callable
```

This appears to be the line raising this exception:

https://github.com/abetlen/llama-cpp-python/blob/99f2ebfde18912adeb7f714b49c1ddb624df3087/llama_cpp/_internals.py#L69

In the case of using the pre-downloaded model, I see the printed model output first before I get the exception. In the case of pulling from Hugging Face Hub, I get the exception immediately.

# Environment and Context

- macOS 15.4 running on M3 Apple Silicon
- Python 3.11.11
- GNU Make 3.81
- llama-cpp-python @ 99f2ebf (the latest from `main` as of 2025-04-11)

# Steps to Reproduce

1. Copy either test script from above into `test-llama.py`.
2. Run `python test-llama.py`

This seems to be an issue specific to the Python bindings, so I did not trying building `llama.cpp`.

This issue is possibly related to #1442.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。