abetlen / abetlen/llama-cpp-python

Retrieve attention score for all input tokens per generated token

Đang mở
#1,141 11 bình luận 1 reaction 0 người được giao Xem trên GitHub
enhancement question
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

**Is your feature request related to a problem? Please describe.**
In RAG-scenarious, I think it would be a great help to differentiate if a LLM is hallucinating or retrieving its informations from the given context, when we could get an attention score for all input-tokens per generated token.

**Describe the solution you'd like**
Having a callback-mechanism for every generated token, similar to the LogitsProcessor, that receives a list of scores.

**Describe alternatives you've considered**
Calculating the scores by myself. But my knowledge of transformers is not sufficient.

**Additional context**
I would like to build something like the "Attention tracing" in [this](https://github.com/mattneary/attention) repository, but with llama.cpp as backend.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.