huggingface / huggingface/diffusers
apple: support for MLX quantized linear in diffusers
- Dominant language
- Python
- Stars
- 34.5k
- Forks
- 7.3k
- Avg merge
- 3d 3h
- Merged PRs (30d)
- 91
Description
**Is your feature request related to a problem? Please describe.**
As an Apple MPS user, it always feels somewhat like we're second-class citizens with respect to the latest and greatest optimisations that only happen for other platforms. The biggest deal is likely xformers/bitsandbytes which remain CUDA-only, but the outcome is more important than the codepath used to get there.
**Describe the solution you'd like.**
I've discovered Apple has some [MLX examples](https://github.com/ml-explore/mlx-examples/blob/main/stable_diffusion/txt2image.py) for T2I inference on SDXL and other SD models that allow AoT quantization of the unet and text encoders.
**Describe alternatives you've considered.**
There is [metal-flash-attention](https://github.com/philipturner/metal-flash-attention) but it would require writing integrating custom Metal kernels, which feels out of scope for Diffusers.
We also have a couple forks of bitsandbytes which aim to improve portability, but there's nothing actionable yet.
**Additional context.**
I haven't tried to implement it yet, it would probably require a bit of monkeying around.
Contributor guide
Assessment
This issue has not been assessed yet.