Comfy-Org / Comfy-Org/comfy-kitchen
RTX30xx FP8 feather matmul triton kernel?
- Dominant language
- Python
- Stars
- 220
- Forks
- 91
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 12
Description
I saw this was released https://github.com/SuriyaaMM/feather not too long ago.
I was wondering, is that something that would ideally be implemented here? Would it basically speed up all fp8 operations presuming the operation was completed long before the memory was copied from VRAM to the registers? Seems like it's in line with the kernels here.
Contributor guide
Research direction
Start by reviewing the existing kernels in this repository and the linked feather project, then determine whether an RTX30xx FP8 matmul Triton kernel fits the library's backends. Done would require a decided implementation scope and evidence about FP8 operation speed; the issue names no files or tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100