abetlen / abetlen/llama-cpp-python

Retrieve attention score for all input tokens per generated token

Abierto
#1,141 11 comentarios 1 reacción 0 asignados Ver en GitHub
enhancement question
Lenguaje dominante
Python
Estrellas
10.6k
Forks
1.4k
Métricas de merge de PR
Métricas de PR pendientes

Descripción

**Is your feature request related to a problem? Please describe.**
In RAG-scenarious, I think it would be a great help to differentiate if a LLM is hallucinating or retrieving its informations from the given context, when we could get an attention score for all input-tokens per generated token.

**Describe the solution you'd like**
Having a callback-mechanism for every generated token, similar to the LogitsProcessor, that receives a list of scores.

**Describe alternatives you've considered**
Calculating the scores by myself. But my knowledge of transformers is not sufficient.

**Additional context**
I would like to build something like the "Attention tracing" in [this](https://github.com/mattneary/attention) repository, but with llama.cpp as backend.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.