huggingface / huggingface/diffusers
better reporting of errors when using TP
- 主要言語
- Python
- スター
- 34.5k
- フォーク
- 7.3k
- 平均マージ
- 3日 3時間
- マージ済み PR(30日)
- 91
説明
I think the following aren't implemented at the moment (which is fine; just flagging).
* **Sharded loading**. from_pretrained should stream shards straight to each rank's DTensor rather than materializing the full checkpoint then slicing — otherwise TP saves you nothing at load time. And check `save_pretrained` / state_dict calls `.full_tensor()` or uses DCP. I think we should at least raise when `save_pretrained()` is called in case TP is enabled?
* **LoRA loading**. for a colwise base layer, lora_A replicated + lora_B colwise; for rowwise, lora_A rowwise + lora_B replicated. If the plan doesn't cover PEFT layers, loading an adapter onto a TP model will either error or be wrong. I think we should detect if the model has `peft` layers injected and raise if TP is requested?
* **Quantization, offloading**. We should probably also raise when these are requested?
_Originally posted by @sayakpaul in https://github.com/huggingface/diffusers/pull/13718#discussion_r3662821938_
コントリビューションガイド
評価
この issue はまだ評価されていません。