huggingface / huggingface/diffusers

better reporting of errors when using TP

オープン
#14,533 コメント 0 件 リアクション 0 件 担当者 1 名 @JingyaHuang が担当を希望しています GitHub で見る
主要言語
Python
スター
34.5k
フォーク
7.3k
平均マージ
3日 3時間
マージ済み PR(30日)
91

説明

I think the following aren't implemented at the moment (which is fine; just flagging).

* **Sharded loading**. from_pretrained should stream shards straight to each rank's DTensor rather than materializing the full checkpoint then slicing — otherwise TP saves you nothing at load time. And check `save_pretrained` / state_dict calls `.full_tensor()` or uses DCP. I think we should at least raise when `save_pretrained()` is called in case TP is enabled?
* **LoRA loading**. for a colwise base layer, lora_A replicated + lora_B colwise; for rowwise, lora_A rowwise + lora_B replicated. If the plan doesn't cover PEFT layers, loading an adapter onto a TP model will either error or be wrong. I think we should detect if the model has `peft` layers injected and raise if TP is requested?
* **Quantization, offloading**. We should probably also raise when these are requested?

_Originally posted by @sayakpaul in https://github.com/huggingface/diffusers/pull/13718#discussion_r3662821938_

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。