huggingface / huggingface/diffusers

apple: support for MLX quantized linear in diffusers

Aperta
#7,675 17 commenti 0 reazioni 0 assegnatari Vedi su GitHub
wip
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

**Is your feature request related to a problem? Please describe.**

As an Apple MPS user, it always feels somewhat like we're second-class citizens with respect to the latest and greatest optimisations that only happen for other platforms. The biggest deal is likely xformers/bitsandbytes which remain CUDA-only, but the outcome is more important than the codepath used to get there.

**Describe the solution you'd like.**

I've discovered Apple has some [MLX examples](https://github.com/ml-explore/mlx-examples/blob/main/stable_diffusion/txt2image.py) for T2I inference on SDXL and other SD models that allow AoT quantization of the unet and text encoders.

**Describe alternatives you've considered.**

There is [metal-flash-attention](https://github.com/philipturner/metal-flash-attention) but it would require writing integrating custom Metal kernels, which feels out of scope for Diffusers.

We also have a couple forks of bitsandbytes which aim to improve portability, but there's nothing actionable yet.

**Additional context.**

I haven't tried to implement it yet, it would probably require a bit of monkeying around.

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.