Project-MONAI / Project-MONAI/MONAI
Enable relative positional embedding in flash attention
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.7k
- Forks
- 1.6k
- Avg merge
- 5d 1h
- Merged PRs (30d)
- 20
Description
From reading this thread:
https://github.com/pytorch/pytorch/issues/96099#issuecomment-1480430583
It seems to me that the relative positional embedding can be integrated with scaled_dot_product_attention 's attn_mask argument. However, it can be slow as it's not taking the "fast path".
Do you think we can keep this option open for users who wants to use flash_attention and rel_pos_embedding?
Originally posted by @mingxin-zheng in https://github.com/Project-MONAI/MONAI/pull/7977#discussion_r1701825032
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the linked PyTorch discussion and the scaled_dot_product_attention attn_mask argument described in the issue. Determine how relative positional embeddings could remain compatible with the flash-attention fast path, then define and verify the expected behavior and performance before changing the integration.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100