huggingface / huggingface/diffusers

Is Lumina2Pipeline's mu calculation correct?

Ouverte
#12,913 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
bug stale
Langage dominant
Python
Étoiles
34.5k
Forks
7.3k
Merge moyen
3 j 3 h
PR mergées (30 j)
91

Description

### Describe the bug

Description

While reviewing the current main-branch implementation of pipeline_lumina2, I noticed a potential bug in the calculation of mu within the pipeline's __call__.

In the following section of the code:

https://github.com/huggingface/diffusers/blob/5ffb65803d0ddc5e3298c35df638ceed5e580922/src/diffusers/pipelines/lumina2/pipeline_lumina2.py#L484-L503

The latent tensor appears to have the shape:

(batch_size, num_channels_latents, height, width)

However, later in the same file:

https://github.com/huggingface/diffusers/blob/5ffb65803d0ddc5e3298c35df638ceed5e580922/src/diffusers/pipelines/lumina2/pipeline_lumina2.py#L699-L706

the value latent.shape[1] (i.e., num_channels_latents) is passed as the argument for image_seq_len when computing mu.
This seems incorrect, since image_seq_len should represent the number of image tokens or sequence length, not the number of latent channels.

Expected Behavior

image_seq_len should likely correspond to the number of spatial tokens derived from (height, width) (or another tokenization step), rather than the number of latent channels.

Actual Behavior

The current implementation uses latent.shape[1] as image_seq_len, which likely leads to unintended behavior in the computation of mu and subsequent sampling steps.

Suggested Fix

Review the logic where image_seq_len is passed, and ensure it reflects the correct sequence length dimension (possibly derived from spatial resolution or token count, rather than channel count).

### Reproduction

At the moment, I don’t have a copy/paste runnable MRE because this was identified via manual logic review rather than reproducing the behavior in a runtime environment.

### Logs

```shell

```

### System Info

Diffusers==0.36.0
Python==3.13

### Who can help?

_No response_

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez dans src/diffusers/pipelines/lumina2/pipeline_lumina2.py en lisant les sections de __call__ autour des lignes 484-503 et 699-706. Suivez la forme du latent et la valeur fournie comme image_seq_len, puis vérifiez que mu utilise la longueur prévue du token d’image ou de la séquence, et ajoutez ou mettez à jour la couverture pour le comportement corrigé.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
machine-learning
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.