OpenImagingLab / OpenImagingLab/FlashVSR
About topk in sparse attention
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.9k
- Forks
- 152
- PR merge metrics
- No merged PRs in 30d
Description
Why do the topk selection across all query-key pairs together but not topk selection across the row for every query respectively. There may be only a part of tokens are updated in the attention module.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read diffsynth/models/wan_video_dit.py around lines 140-149 and inspect how topk is applied across query-key pairs. Compare global selection with per-query row selection, then determine whether the attention module intentionally updates only part of the tokens. Done means the intended behavior and its rationale are documented or the required change is clearly specified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100