deepseek-ai / deepseek-ai/FlashMLA
Ampere architecture FlashMLA bring-up
- Dominant language
- C++
- Stars
- 12.9k
- Forks
- 1.2k
- Avg merge
- 4h 20m
- Merged PRs (30d)
- 2
Description
Based on the Hopper architecture FlashMLA and FlashAttentionV2, the FlashMLA based on the Ampere architecture is implemented.
Without using double buffer optimization, the performance can achiev up to `300 GB/s` in and `125 TFLOPS` on A100.
This is the code repository:
https://github.com/pzhao-eng/FlashMLA
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the issue description and the linked pzhao-eng/FlashMLA repository, then compare its Ampere implementation with this repository's Hopper FlashMLA and FlashAttentionV2 code. The issue does not name files, tests, or an integration target; completion criteria should be clarified before work begins.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100