Comfy-Org / Comfy-Org/comfy-kitchen

FP8 E4M3 quantized models inference is slower than same model quantized in MXFP8

Open
#38 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
220
Forks
91
Avg merge
1d 7h
Merged PRs (30d)
12

Description

It's quite weird. I think mxfp8 inference is a bit more complex than the legacy fp8e4m3. So does it mean we still have some potential in legacy fp8 inference?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.