alipay / alipay/PainlessInferenceAcceleration

modeling_qwen attention not use multi branch position ids & attention_mask

Open
#27 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
370
Forks
23
PR merge metrics
No merged PRs in 30d

Description

I reviewed the code of modeling_qwen.py, and I noticed that, within the lookahead process, the draft_ids matched from the TrieTree are such that the attention_mask and position ids associated with these draft_ids are not being utilized in the attention mechanism. This, I believe, might be an implementation error. Could you please point out where my understanding is incorrect?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.