huggingface / huggingface/diffusers
apple: support for MLX quantized linear in diffusers
- 主要言語
- Python
- スター
- 34.5k
- フォーク
- 7.3k
- 平均マージ
- 3日 3時間
- マージ済み PR(30日)
- 91
説明
**Is your feature request related to a problem? Please describe.**
As an Apple MPS user, it always feels somewhat like we're second-class citizens with respect to the latest and greatest optimisations that only happen for other platforms. The biggest deal is likely xformers/bitsandbytes which remain CUDA-only, but the outcome is more important than the codepath used to get there.
**Describe the solution you'd like.**
I've discovered Apple has some [MLX examples](https://github.com/ml-explore/mlx-examples/blob/main/stable_diffusion/txt2image.py) for T2I inference on SDXL and other SD models that allow AoT quantization of the unet and text encoders.
**Describe alternatives you've considered.**
There is [metal-flash-attention](https://github.com/philipturner/metal-flash-attention) but it would require writing integrating custom Metal kernels, which feels out of scope for Diffusers.
We also have a couple forks of bitsandbytes which aim to improve portability, but there's nothing actionable yet.
**Additional context.**
I haven't tried to implement it yet, it would probably require a bit of monkeying around.
コントリビューションガイド
調査の方向性
Start with the linked MLX example, stable_diffusion/txt2image.py, and review how Diffusers currently handles quantized linear layers and Apple MPS inference. Determine the integration scope for SDXL and other Stable Diffusion models; done means MLX quantized linear inference works on Apple MPS with documented coverage and validation.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python, pytorch
- 領域
- machine-learning, performance
- issue の種類
- 機能追加
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 25/100