huggingface / huggingface/diffusers

[SD3] Incorrect stochastic sampling implementation

Open
#8,549 5 comments 0 reactions 0 assignees View on GitHub
bug stale
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

### Describe the bug

Ref to Algorithm 2 of [EDM](https://arxiv.org/abs/2206.00364), for a given sample $x_t$, noise is introduced to it reaching a higher noise level $\hat{t}$, then we evaluate network with $\hat{x_t}$, $\hat{t}$ as input. However, the current implementation evaluates network with $x_t$, $t$ as input, which is inconsistent from definition.

Current implementation is more similar with Euler-Maruyama in spirit, "One can interpret Euler–Maruyama as first adding
noise and then performing an ODE step, not from the intermediate state after noise injection, but
assuming that $x$ and $\sigma$ remained at the initial state at the beginning of the iteration step." quote from [EDM](https://arxiv.org/abs/2206.00364)

### Reproduction

no

### Logs

_No response_

### System Info

no

### Who can help?

@yiyixuxu @sayakpaul

Contributor guide

Open the contributing guide

Research direction

No file, test, or entry point is named. Start by locating the stochastic sampling implementation and compare its network inputs with Algorithm 2 of the linked EDM paper, focusing on whether the noise-adjusted sample and noise level are evaluated. Done means the implementation follows that algorithm and has coverage demonstrating the corrected inputs.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.