alipay / alipay/PainlessInferenceAcceleration

modeling_qwen attention not use multi branch position ids & attention_mask

未关闭
#27 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
370
派生
23
PR 合并指标
30 天内没有已合并 PR

描述

I reviewed the code of modeling_qwen.py, and I noticed that, within the lookahead process, the draft_ids matched from the TrieTree are such that the attention_mask and position ids associated with these draft_ids are not being utilized in the attention mechanism. This, I believe, might be an implementation error. Could you please point out where my understanding is incorrect?

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。