deepseek-ai / deepseek-ai/FlashMLA

Ampere architecture FlashMLA bring-up

Open
#66 2 comments 9 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
12.9k
Forks
1.2k
Avg merge
4h 20m
Merged PRs (30d)
2

Description

Based on the Hopper architecture FlashMLA and FlashAttentionV2, the FlashMLA based on the Ampere architecture is implemented.
Without using double buffer optimization, the performance can achiev up to `300 GB/s` in and `125 TFLOPS` on A100.
This is the code repository:
https://github.com/pzhao-eng/FlashMLA

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the issue description and the linked pzhao-eng/FlashMLA repository, then compare its Ampere implementation with this repository's Hopper FlashMLA and FlashAttentionV2 code. The issue does not name files, tests, or an integration target; completion criteria should be clarified before work begins.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
machine-learning, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.