abetlen / abetlen/llama-cpp-python

Can't compute multiple embeddings in a single call

未关闭
#2,051 4 条评论 5 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
10.6k
派生
1.4k
平均合并
5 小时 23 分钟
30 天内合并 PR
5

描述

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [x] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Running this code:

```python
model = llama_cpp.Llama ("mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
embeddings = model.embed (["Hello", "World"])
```

used to work in v0.3.14

# Current Behavior

The code raises an exception `RuntimeError: llama_decode returned -1`. The following messages are printed to the console:

```
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
```

# Environment and Context

llama-cpp-python was compiled in CUDA mode

# Failure Information (for bugs)

Please help provide information about the failure if this is a bug. If it is not a bug, please remove the rest of this template.

# Steps to Reproduce

```python
Python 3.11.2 (main, Apr 28 2025, 14:11:48) [GCC 12.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import llama_cpp
>>> model = llama_cpp.Llama ("../models/mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
...
>>> embeddings = model.embed (["Hello", "World"])
decode: cannot decode batches with this context (calling encode() instead)
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
llama_decode: failed to decode, ret = -1
Traceback (most recent call last):
File "", line 1, in
File ".../site-packages/llama_cpp/llama.py", line 1108, in embed
decode_batch(s_batch)
File ".../site-packages/llama_cpp/llama.py", line 1045, in decode_batch
self._ctx.decode(self._batch)
File ".../site-packages/llama_cpp/_internals.py", line 327, in decode
raise RuntimeError(f"llama_decode returned {return_code}")
RuntimeError: llama_decode returned -1
```

贡献指南

打开贡献指南

调研方向

错误发生在 llama_cpp/llama.py 中约第 1108 行的 embed 方法以及第 1045 行的 decode_batch 中。首先检查 C++ 绑定(_internals.py)中的 batch 初始化和序列 ID 处理。查看 v0.3.14 与当前版本之间 embedding 逻辑的变化。运行提供的复现脚本以查看确切的错误,然后检查 llama.cpp 库对多个序列进行 batch 解码的实现。

由索引模型根据 Issue 内容生成。

评估

技术栈
c, python
领域
ai, machine-learning
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
描述清楚
新手友好度
40/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。