deepspeedai / deepspeedai/DeepSpeed
Question about the definitions in block sparse attention
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Hi, I have some question regarding the block sparse attention.
If I understand the description of API correctly, block is the block size (i.e., number of tokens in a block) while num_local_blocks denotes the number of blocks (#tokens_per_window = block * num_local_blocks) in a local window. So no matter which value (unidirectional or bidirectional) I choose for attention, the tokens within a block will attend each other?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file, test, or entry point is named. Start by locating the block sparse attention API documentation and implementation, then verify the meanings of block, num_local_blocks, and attention. Done means the documentation clearly answers whether tokens within a block attend each other under both attention modes.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100