huggingface / huggingface/diffusers

[Quantization] Add support for Comfy Quants backend

Open
#14,705 3 comments 0 reactions 1 assignee Claimed by @PrakshaaleJain View on GitHub
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

Add support for the [Comfy Quant](https://github.com/Comfy-Org/comfy-quants) quantization toolkit

Supported formats: FP8 (E4M3/E5M2), INT8 (W8A8, tensorwise), MXFP8, NVFP4, INT4 (SVDQuant W4A4, AWQ W4A16).

comfy-quants is export-only. The inference needs to run with [`comfy-kitchen`](https://github.com/Comfy-Org/comfy-kitchen)

### Proposed approach

Add a `ComfyQuantConfig` / `ComfyQuantizer` that uses `comfy-kitchen` to wrap weights as a `QuantizedTensor` with the appropriate layout.

```python
from diffusers import FluxTransformer2DModel, ComfyQuantConfig

config = ComfyQuantConfig(compute_dtype=torch.bfloat16)
model = FluxTransformer2DModel.from_single_file(
"path/to/comfy_quant_checkpoint.safetensors",
quantization_config=config,
)
```

### Related
• comfy-quants https://github.com/Comfy-Org/comfy-quants
• comfy-kitchen https://github.com/Comfy-Org/comfy-kitchen

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.