abetlen / abetlen/llama-cpp-python
Retrieve attention score for all input tokens per generated token
- Lingua principale
- Python
- Stelle
- 10.6k
- Fork
- 1.4k
- Metriche di merge delle PR
- Metriche PR in attesa
Descrizione
**Is your feature request related to a problem? Please describe.**
In RAG-scenarious, I think it would be a great help to differentiate if a LLM is hallucinating or retrieving its informations from the given context, when we could get an attention score for all input-tokens per generated token.
**Describe the solution you'd like**
Having a callback-mechanism for every generated token, similar to the LogitsProcessor, that receives a list of scores.
**Describe alternatives you've considered**
Calculating the scores by myself. But my knowledge of transformers is not sufficient.
**Additional context**
I would like to build something like the "Attention tracing" in [this](https://github.com/mattneary/attention) repository, but with llama.cpp as backend.
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.