huggingface / huggingface/diffusers
[SD3 ControlNet] bug in pipeline 'controlnet_pooled_projections'
- 主要言語
- Python
- スター
- 34.5k
- フォーク
- 7.3k
- 平均マージ
- 3日 3時間
- マージ済み PR(30日)
- 91
説明
### Describe the bug
Hi,
I think I found an issue that causes a misalignment between training and inference in SD3 ControlNet.
https://github.com/huggingface/diffusers/blob/a3e8d3f7deed140f57a28d82dd0b5d965bd0fb09/src/diffusers/pipelines/controlnet_sd3/pipeline_stable_diffusion_3_controlnet.py#L977
I think the if-else block starting there is not correct. It should be
```python
if controlnet_pooled_projections is None and pooled_prompt_embeds is None:
controlnet_pooled_projections = torch.zeros_like(pooled_prompt_embeds)
elif controlnet_pooled_projections is None:
controlnet_pooled_projections = pooled_prompt_embeds
```
Given that in training, the pooled_prompt_embeds are fed to the model:
https://github.com/huggingface/diffusers/blob/a3e8d3f7deed140f57a28d82dd0b5d965bd0fb09/examples/controlnet/train_controlnet_sd3.py#L1293
Additionally, I am wondering if this line:
https://github.com/huggingface/diffusers/blob/a3e8d3f7deed140f57a28d82dd0b5d965bd0fb09/examples/controlnet/train_controlnet_sd3.py#L1287
Should be aligned with this line:
https://github.com/huggingface/diffusers/blob/a3e8d3f7deed140f57a28d82dd0b5d965bd0fb09/examples/controlnet/train_controlnet_sd3.py#L1257
This seems to be the more sensible approach, but will probably not make much difference since the ControlNet can also learn the shift. It might speed up convergence *slightly*.
Best,
Tobias
### Reproduction
Train an SD3 ControlNet and during log_validation it will be executed.
### Logs
_No response_
### System Info
diffusers==0.30.3
### Who can help?
@yiyixuxu @sayakpaul
コントリビューションガイド
調査の方向性
Start with the if-else block around line 977 in src/diffusers/pipelines/controlnet_sd3/pipeline_stable_diffusion_3_controlnet.py, then compare it with the pooled prompt handling around lines 1287 and 1293 in examples/controlnet/train_controlnet_sd3.py. Run the stated log_validation path after training and verify that training and inference use aligned pooled projections, including a resolved decision about the shift at line 1257.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python, pytorch
- 領域
- machine-learning
- issue の種類
- バグ
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 活発さ
- 停滞
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 45/100