huggingface / huggingface/diffusers

training example for instruct pix2pix doesn't zero out embeds

オープン
#7,920 コメント 9 件 リアクション 0 件 担当者 0 名 GitHub で見る
bug stale
主要言語
Python
スター
34.5k
フォーク
7.3k
平均マージ
3日 3時間
マージ済み PR(30日)
91

説明

### Describe the bug

When running inference on SDXL, the config specifies to zero out the embedding when the prompt is empty.

### Reproduction

```py
# Get null conditioning
def compute_null_conditioning():
null_conditioning_list = []
for a_tokenizer, a_text_encoder in zip(tokenizers, text_encoders):
null_conditioning_list.append(
a_text_encoder(
tokenize_captions([""], tokenizer=a_tokenizer).to(accelerator.device),
output_hidden_states=True,
).hidden_states[-2]
)
return torch.concat(null_conditioning_list, dim=-1)

null_conditioning = compute_null_conditioning()
```

this could likely be replaced with a probabilistic call to `torch.zeros_like()` inside the training loop instead.

I've checked the values of the embeds, and classifier-free guidance at inference time definitely makes use of the zero embed and not just `""`, which end up producing very different results.

other models though like deepfloyd just use `""` from eg. T5 and behave rather differently.

### Logs

_No response_

### System Info

N/A

### Who can help?

@sayakpaul

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

Start in the instruct pix2pix training example, inspect compute_null_conditioning and the training loop, and compare their conditioning behavior with the SDXL inference configuration. Done means the example handles empty-prompt conditioning consistently with the configured zero embeddings without changing the behavior required by other model families.

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python, pytorch
領域
machine-learning
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。