huggingface / huggingface/diffusers

question about training diffusion-inpainting model

オープン
#6,502 コメント 5 件 リアクション 2 件 担当者 0 名 GitHub で見る
stale
主要言語
Python
スター
34.5k
フォーク
7.3k
平均マージ
3日 3時間
マージ済み PR(30日)
91

説明

Hi, everyone! I'm struggling with the inpainting/outpainting, and quite confused about the input of the model in the training stage, hope can get some help😢

In **diffusers/examples/research_projects/multi_subject_dreambooth_inpainting/train_multi_subject_dreambooth_inpainting.py** @gzguevara and **diffusers/examples/research_projects/dreambooth_inpaint/train_dreambooth_inpaint.py**, the input of 9-ch inpainting model are the combination of gt(add noise), masked_img, and mask during training.
|gt imgs|masked imgs|mask|
|--------|---------------|------|
|![img](https://github.com/huggingface/diffusers/assets/50061868/70ecf472-7ac0-40e1-88b1-b5592da7d360)|![masked_img](https://github.com/huggingface/diffusers/assets/50061868/14c5e53c-b415-40c2-91e4-9132a3bb72b4)|![msk](https://github.com/huggingface/diffusers/assets/50061868/bcb4fff3-ccf4-4c27-90b5-dd7e3f76faba)|

I am curious about why GT image can be input into the unet directly. Even though it has been added with noise, it is still visible to the unet.
```
latents = vae.encode(batch["pixel_values"].to(dtype=weight_dtype)).latent_dist.sample()
latents = latents * vae.config.scaling_factor

masked_latents = vae.encode(batch["masked_images"].reshape(batch["pixel_values"].shape).to(dtype=weight_dtype)).latent_dist.sample()
masked_latents = masked_latents * vae.config.scaling_factor

masks = batch["masks"]
mask = torch.stack([torch.nn.functional.interpolate(mask, size=(args.resolution // 8, args.resolution // 8)) for mask in masks])
mask = mask.reshape(-1, 1, args.resolution // 8, args.resolution // 8)

noise = torch.randn_like(latents)
bsz = latents.shape[0]
timesteps = torch.randint(0, noise_scheduler.config.num_train_timesteps, (bsz,), device=latents.device)
timesteps = timesteps.long()
noisy_latents = noise_scheduler.add_noise(latents, noise, timesteps)

latent_model_input = torch.cat([noisy_latents, mask, masked_latents], dim=1)
```
Use the images above as an example: the input is car image, and the expected output is car image during training. And when comes for infering, users can add objects to the image(eg. the input is image unrelated to car(arbitrary object or just background), and the expected output is car image.) There is a gap between training and testing.

I think it may benefits from text-guided effect, but I still have doubts. On the one hand, model needs GT to be optimized, and it is often used as a target in other generative model, rather than as a direct input to the model. On the other hand, diffusion model predict Gaussian noise, there seems to be no other way for diffusion model to be constrained from gt.

When turns to outpainting task, the gap between training and testing is bigger: If I use the combination of gt(add noise), masked_img, and mask for training, waht should I pad the image with unmasked area for infering.I don't understand how does the model avoid learning a simple mapping, I'd be grateful if anyone could give me advice.

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

Begin with the two training scripts named in the issue, then compare their latent construction with the corresponding inference path. A useful resolution should explain whether the training inputs and inference inputs are aligned for inpainting and outpainting, including what fills the unmasked area.

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python, pytorch
領域
computer-vision, machine-learning
issue の種類
ドキュメント
難易度
5/5
見積もり時間
1週間以上
活発さ
静か
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。