deepspeedai / deepspeedai/DeepSpeed
[REQUEST] [TRITON] Upgrade Sparse Attention by Using Triton > 2.1
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Hi, we are looking at deepspeed.ops.sparse_attention and find out that current SA is based on triton==1.0.0, which is old version. Current triton is 2.x and our supported version is 2.x. May I know if there is any plan on upgrading triton version to 2.x and maintain the sparse attention kernel?
My error stack is mainly on deepspeed.ops.sparse_attention.matmul:
import triton._C.libtriton as libtriton
segmented = libtriton.superblock(layout.data_ptr(),
layout.shape[0],
layout.shape[1],
layout.shape[2],
start_width)
Thanks!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with deepspeed/ops/sparse_attention/matmul.py, especially the libtriton.superblock call shown in the issue, and inspect how the current sparse attention implementation depends on Triton 1.0.0. Determine the compatibility changes needed for supported Triton 2.x, then verify that the sparse attention kernel is maintained and the reported error no longer occurs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100