huggingface / huggingface/diffusers

deepfloyd stage 2 crashes with tensor size mismatch when input image size is not divisible by 8

Aperta
#7,842 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub
bug stale
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

### Describe the bug

DeepFloyd's upstream code supports 8px-aligned inputs for stage II, which I believe the Diffusers implementation is based upon. However, it seems that for certain sizes, there is some unfortunate interaction between the hidden states and the residual hidden states.

I'm not sure if this is something fundamental to the model - if it is, we probably want to understand the conditions under which this problem occurs and provide an error to the user about an incompatible resolution.

### Reproduction

```py
from diffusers import IFSuperResolutionPipeline
import torch
from PIL import Image
import numpy as np

torch.manual_seed(42)

# Configuration for initial image and desired output
initial_width = 86 # Adjusted width to be one-fourth of 344 (approximately)
initial_height = 64 # Adjusted height to be one-fourth of 256

# Initialize your device setting based on availability
torch_device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "xpu" if torch.xpu.is_available() else "cpu"

# Create a dummy image (86x64)
dummy_image = torch.rand((3, initial_height, initial_width), dtype=torch.float32) # Random noise image
dummy_image = (dummy_image * 255).to(torch.uint8) # Convert to 8-bit format
dummy_pil_image = Image.fromarray(dummy_image.numpy().transpose(1, 2, 0)) # Convert to PIL image for compatibility
dummy_pil_image.save("dummy_input.png") # Save the initial dummy image

# Load your stage 2 pipeline
stage2_pipe = IFSuperResolutionPipeline.from_pretrained("DeepFloyd/IF-II-M-v1.0", watermarker=None, safety_checker=None, local_files_only=False).to(device=torch_device, dtype=torch.bfloat16)

# Upscale the dummy image using stage 2 of the pipeline
upscaled_image = stage2_pipe(
prompt="A simple upscaled image",
image=dummy_pil_image,
guidance_scale=5.5,
num_inference_steps=20,
width=344,
height=256
).images[0]

upscaled_image.save("upscaled_dummy_output.png")
```

### Logs

```shell
0%| | 0/20 [00:00

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia dal punto di ingresso IFSuperResolutionPipeline e riproduci l’esecuzione dello stage II descritta nell’issue con l’input 86x64 e l’output 344x256. Traccia le shape di hidden_states e res_hidden_states mostrate nei log, determina se la risoluzione debba essere rifiutata o gestita e verifica che la riproduzione non vada più in crash o segnali un errore di incompatibilità appropriato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python, pytorch
Ambito
computer-vision, machine-learning
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.