huggingface / huggingface/diffusers
apple: support for MLX quantized linear in diffusers
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
**Is your feature request related to a problem? Please describe.**
As an Apple MPS user, it always feels somewhat like we're second-class citizens with respect to the latest and greatest optimisations that only happen for other platforms. The biggest deal is likely xformers/bitsandbytes which remain CUDA-only, but the outcome is more important than the codepath used to get there.
**Describe the solution you'd like.**
I've discovered Apple has some [MLX examples](https://github.com/ml-explore/mlx-examples/blob/main/stable_diffusion/txt2image.py) for T2I inference on SDXL and other SD models that allow AoT quantization of the unet and text encoders.
**Describe alternatives you've considered.**
There is [metal-flash-attention](https://github.com/philipturner/metal-flash-attention) but it would require writing integrating custom Metal kernels, which feels out of scope for Diffusers.
We also have a couple forks of bitsandbytes which aim to improve portability, but there's nothing actionable yet.
**Additional context.**
I haven't tried to implement it yet, it would probably require a bit of monkeying around.
Beitragsleitfaden
Rechercherichtung
Start with the linked MLX example, stable_diffusion/txt2image.py, and review how Diffusers currently handles quantized linear layers and Apple MPS inference. Determine the integration scope for SDXL and other Stable Diffusion models; done means MLX quantized linear inference works on Apple MPS with documented coverage and validation.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, pytorch
- Bereich
- machine-learning, performance
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100