abetlen / abetlen/llama-cpp-python

Retrieve attention score for all input tokens per generated token

未關閉
#1,141 11 則留言 1 個 reaction 已指派 0 人 在 GitHub 檢視
enhancement question
主要語言
Python
星號
10.6k
分支
1.4k
PR 合併指標
PR 指標待擷取

描述

**Is your feature request related to a problem? Please describe.**
In RAG-scenarious, I think it would be a great help to differentiate if a LLM is hallucinating or retrieving its informations from the given context, when we could get an attention score for all input-tokens per generated token.

**Describe the solution you'd like**
Having a callback-mechanism for every generated token, similar to the LogitsProcessor, that receives a list of scores.

**Describe alternatives you've considered**
Calculating the scores by myself. But my knowledge of transformers is not sufficient.

**Additional context**
I would like to build something like the "Attention tracing" in [this](https://github.com/mattneary/attention) repository, but with llama.cpp as backend.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。