huggingface / huggingface/diffusers
training example for instruct pix2pix doesn't zero out embeds
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
### Describe the bug
When running inference on SDXL, the config specifies to zero out the embedding when the prompt is empty.
### Reproduction
```py
# Get null conditioning
def compute_null_conditioning():
null_conditioning_list = []
for a_tokenizer, a_text_encoder in zip(tokenizers, text_encoders):
null_conditioning_list.append(
a_text_encoder(
tokenize_captions([""], tokenizer=a_tokenizer).to(accelerator.device),
output_hidden_states=True,
).hidden_states[-2]
)
return torch.concat(null_conditioning_list, dim=-1)
null_conditioning = compute_null_conditioning()
```
this could likely be replaced with a probabilistic call to `torch.zeros_like()` inside the training loop instead.
I've checked the values of the embeds, and classifier-free guidance at inference time definitely makes use of the zero embed and not just `""`, which end up producing very different results.
other models though like deepfloyd just use `""` from eg. T5 and behave rather differently.
### Logs
_No response_
### System Info
N/A
### Who can help?
@sayakpaul
Beitragsleitfaden
Rechercherichtung
Start in the instruct pix2pix training example, inspect compute_null_conditioning and the training loop, and compare their conditioning behavior with the SDXL inference configuration. Done means the example handles empty-prompt conditioning consistently with the configured zero embeddings without changing the behavior required by other model families.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, pytorch
- Bereich
- machine-learning
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 35/100