Comfy-Org / Comfy-Org/comfy-kitchen

RTX30xx FP8 feather matmul triton kernel?

Open
#16 6 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
220
Forks
91
Avg merge
1d 7h
Merged PRs (30d)
12

Description

I saw this was released https://github.com/SuriyaaMM/feather not too long ago.

I was wondering, is that something that would ideally be implemented here? Would it basically speed up all fp8 operations presuming the operation was completed long before the memory was copied from VRAM to the registers? Seems like it's in line with the kernels here.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the existing kernels in this repository and the linked feather project, then determine whether an RTX30xx FP8 matmul Triton kernel fits the library's backends. Done would require a decided implementation scope and evidence about FP8 operation speed; the issue names no files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.