deepseek-ai / deepseek-ai/DeepGEMM
Support for Ada Lovelace & Blackwell
Open
- Dominant language
- Cuda
- Stars
- 7.8k
- Forks
- 1.3k
- Avg merge
- 3d 7h
- Merged PRs (30d)
- 3
Description
This is perhaps better served as a feature request, but I can see wide interest in proper F8 implementation for the 4090/5090 series of GPUs. Is there any interest in adapting this for sm_89 & sm_100/a? Is it a steep challenge or is it feasible?
Many thanks for providing this repo. It is a remarkable contribution to the field.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.