deepseek-ai / deepseek-ai/DeepGEMM

Support for Ada Lovelace & Blackwell

Open
#6 13 comments 10 reactions 0 assignees View on GitHub
Dominant language
Cuda
Stars
7.8k
Forks
1.3k
Avg merge
3d 7h
Merged PRs (30d)
3

Description

This is perhaps better served as a feature request, but I can see wide interest in proper F8 implementation for the 4090/5090 series of GPUs. Is there any interest in adapting this for sm_89 & sm_100/a? Is it a steep challenge or is it feasible?

Many thanks for providing this repo. It is a remarkable contribution to the field.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.