abetlen / abetlen/llama-cpp-python

Can't compute multiple embeddings in a single call

Open
#2,051 4 comments 5 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10.6k
Forks
1.4k
Avg merge
5h 23m
Merged PRs (30d)
5

Description

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [x] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Running this code:

```python
model = llama_cpp.Llama ("mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
embeddings = model.embed (["Hello", "World"])
```

used to work in v0.3.14

# Current Behavior

The code raises an exception `RuntimeError: llama_decode returned -1`. The following messages are printed to the console:

```
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
```

# Environment and Context

llama-cpp-python was compiled in CUDA mode

# Failure Information (for bugs)

Please help provide information about the failure if this is a bug. If it is not a bug, please remove the rest of this template.

# Steps to Reproduce

```python
Python 3.11.2 (main, Apr 28 2025, 14:11:48) [GCC 12.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import llama_cpp
>>> model = llama_cpp.Llama ("../models/mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
...
>>> embeddings = model.embed (["Hello", "World"])
decode: cannot decode batches with this context (calling encode() instead)
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
llama_decode: failed to decode, ret = -1
Traceback (most recent call last):
File "", line 1, in
File ".../site-packages/llama_cpp/llama.py", line 1108, in embed
decode_batch(s_batch)
File ".../site-packages/llama_cpp/llama.py", line 1045, in decode_batch
self._ctx.decode(self._batch)
File ".../site-packages/llama_cpp/_internals.py", line 327, in decode
raise RuntimeError(f"llama_decode returned {return_code}")
RuntimeError: llama_decode returned -1
```

Contributor guide

Open the contributing guide

Research direction

The error occurs in llama_cpp/llama.py around line 1108 in the embed method and line 1045 in decode_batch. Start by examining the batch initialization and sequence ID handling in the C++ bindings (_internals.py). Look at the embedding logic changes between v0.3.14 and the current version. Run the provided reproduction script to see the exact error, then check the llama.cpp library's batch decoding for multiple sequences.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, python
Domain
ai, machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
40/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.