deepseek-ai / deepseek-ai/FlashMLA
Built FlashMLA for Windows and Nvidia sm_120 (workstation/50s) cards
- Dominant language
- C++
- Stars
- 12.9k
- Forks
- 1.2k
- Avg merge
- 4h 20m
- Merged PRs (30d)
- 2
Description
I will PR this into this repo at some point, but I think I got a working source for windows MSVC with CUTLASS 4.1 and CUDA 12.9
For now I will just leave this here for crawlers and indexers, have not bench-marked (fully) anything. Will get results at some point.
FYI. at the end I only got sm_120(ONLY) to build so sm_100a will not have support as with sm_90 (I hate Nvidia sm) (Thanks for attending my Ted talk)
https://github.com/IISuperluminaLII/FlashMLA_Windows_sm120
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the linked FlashMLA_Windows_sm120 repository and compare its Windows MSVC source with this repository. Review the stated CUTLASS 4.1, CUDA 12.9, and sm_120-only constraints, then determine the required integration and benchmarking work; the issue does not define completion criteria.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- build-system, performance
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100