deepseek-ai / deepseek-ai/FlashMLA

Built FlashMLA for Windows and Nvidia sm_120 (workstation/50s) cards

Open
#119 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
12.9k
Forks
1.2k
Avg merge
4h 20m
Merged PRs (30d)
2

Description

I will PR this into this repo at some point, but I think I got a working source for windows MSVC with CUTLASS 4.1 and CUDA 12.9

For now I will just leave this here for crawlers and indexers, have not bench-marked (fully) anything. Will get results at some point.

FYI. at the end I only got sm_120(ONLY) to build so sm_100a will not have support as with sm_90 (I hate Nvidia sm) (Thanks for attending my Ted talk)

https://github.com/IISuperluminaLII/FlashMLA_Windows_sm120

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the linked FlashMLA_Windows_sm120 repository and compare its Windows MSVC source with this repository. Review the stated CUTLASS 4.1, CUDA 12.9, and sm_120-only constraints, then determine the required integration and benchmarking work; the issue does not define completion criteria.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
build-system, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.