huggingface / huggingface/diffusers
apple: support for MLX quantized linear in diffusers
- Lenguaje dominante
- Python
- Estrellas
- 34.5k
- Forks
- 7.3k
- Merge medio
- 3 d 3 h
- PR fusionados (30 d)
- 91
Descripción
**Is your feature request related to a problem? Please describe.**
As an Apple MPS user, it always feels somewhat like we're second-class citizens with respect to the latest and greatest optimisations that only happen for other platforms. The biggest deal is likely xformers/bitsandbytes which remain CUDA-only, but the outcome is more important than the codepath used to get there.
**Describe the solution you'd like.**
I've discovered Apple has some [MLX examples](https://github.com/ml-explore/mlx-examples/blob/main/stable_diffusion/txt2image.py) for T2I inference on SDXL and other SD models that allow AoT quantization of the unet and text encoders.
**Describe alternatives you've considered.**
There is [metal-flash-attention](https://github.com/philipturner/metal-flash-attention) but it would require writing integrating custom Metal kernels, which feels out of scope for Diffusers.
We also have a couple forks of bitsandbytes which aim to improve portability, but there's nothing actionable yet.
**Additional context.**
I haven't tried to implement it yet, it would probably require a bit of monkeying around.
Guía de contribución
Línea de trabajo
Start with the linked MLX example, stable_diffusion/txt2image.py, and review how Diffusers currently handles quantized linear layers and Apple MPS inference. Determine the integration scope for SDXL and other Stable Diffusion models; done means MLX quantized linear inference works on Apple MPS with documented coverage and validation.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python, pytorch
- Área
- machine-learning, performance
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100