huggingface / huggingface/diffusers

training example for instruct pix2pix doesn't zero out embeds

Open
#7,920 9 comments 0 reactions 0 assignees View on GitHub
bug stale
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

### Describe the bug

When running inference on SDXL, the config specifies to zero out the embedding when the prompt is empty.

### Reproduction

```py
# Get null conditioning
def compute_null_conditioning():
null_conditioning_list = []
for a_tokenizer, a_text_encoder in zip(tokenizers, text_encoders):
null_conditioning_list.append(
a_text_encoder(
tokenize_captions([""], tokenizer=a_tokenizer).to(accelerator.device),
output_hidden_states=True,
).hidden_states[-2]
)
return torch.concat(null_conditioning_list, dim=-1)

null_conditioning = compute_null_conditioning()
```

this could likely be replaced with a probabilistic call to `torch.zeros_like()` inside the training loop instead.

I've checked the values of the embeds, and classifier-free guidance at inference time definitely makes use of the zero embed and not just `""`, which end up producing very different results.

other models though like deepfloyd just use `""` from eg. T5 and behave rather differently.

### Logs

_No response_

### System Info

N/A

### Who can help?

@sayakpaul

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.