abetlen / abetlen/llama-cpp-python
Can't compute multiple embeddings in a single call
- Vorherrschende Sprache
- Python
- Sterne
- 10.6k
- Forks
- 1.4k
- Ø Merge
- 5 Std. 23 Min.
- Gemergte PRs (30 T.)
- 5
Beschreibung
# Prerequisites
Please answer the following questions for yourself before submitting an issue.
- [x] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.
# Expected Behavior
Running this code:
```python
model = llama_cpp.Llama ("mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
embeddings = model.embed (["Hello", "World"])
```
used to work in v0.3.14
# Current Behavior
The code raises an exception `RuntimeError: llama_decode returned -1`. The following messages are printed to the console:
```
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
```
# Environment and Context
llama-cpp-python was compiled in CUDA mode
# Failure Information (for bugs)
Please help provide information about the failure if this is a bug. If it is not a bug, please remove the rest of this template.
# Steps to Reproduce
```python
Python 3.11.2 (main, Apr 28 2025, 14:11:48) [GCC 12.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import llama_cpp
>>> model = llama_cpp.Llama ("../models/mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
...
>>> embeddings = model.embed (["Hello", "World"])
decode: cannot decode batches with this context (calling encode() instead)
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
llama_decode: failed to decode, ret = -1
Traceback (most recent call last):
File "", line 1, in
File ".../site-packages/llama_cpp/llama.py", line 1108, in embed
decode_batch(s_batch)
File ".../site-packages/llama_cpp/llama.py", line 1045, in decode_batch
self._ctx.decode(self._batch)
File ".../site-packages/llama_cpp/_internals.py", line 327, in decode
raise RuntimeError(f"llama_decode returned {return_code}")
RuntimeError: llama_decode returned -1
```
Beitragsleitfaden
Rechercherichtung
Der Fehler tritt in llama_cpp/llama.py ungefähr bei Zeile 1108 in der embed-Methode und bei Zeile 1045 in decode_batch auf. Untersuche zunächst die Batch-Initialisierung und die Handhabung der Sequenz-IDs in den C++-Bindings (_internals.py). Sieh dir die Änderungen an der Embedding-Logik zwischen v0.3.14 und der aktuellen Version an. Führe das bereitgestellte Reproduktionsskript aus, um den genauen Fehler zu sehen, und überprüfe anschließend die Batch-Dekodierung der llama.cpp-Bibliothek für mehrere Sequenzen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- c, python
- Bereich
- ai, machine-learning
- Issue-Typ
- Bug
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Klar beschrieben
- Anfängerfreundlichkeit
- 40/100