abetlen / abetlen/llama-cpp-python

tensor read out of bounds

Aberta
#1,241 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
10.6k
Forks
1.4k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [x] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [x] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [x] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [x] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Expected Behavior

Stream responses correctly, or, if there is an error, throw an error but not crash the application

# Current Behavior

I'm using it in a Flask app and when multiple generations are running concurrently I get `tensor read out of bounds`

# Environment and Context

Latest version of macOS

# Failure Logs

```
GGML_ASSERT: /private/var/folders/n4/********/llama-cpp-python_**********/vendor/llama.cpp/ggml-backend.c:206: offset + size <= ggml_nbytes(tensor) && "tensor read out of bounds"
```

Guia de contribuição

Abrir o guia de contribuição

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.