huggingface / huggingface/diffusers

Freeing GPU memory after `torch.compile` StableDiffusionXLPipeline UNet

Abierto
#9,530 4 comentarios 0 reacciones 0 asignados Ver en GitHub
stale torch-compile
Lenguaje dominante
Python
Estrellas
34.5k
Forks
7.3k
Merge medio
3 d 3 h
PR fusionados (30 d)
91

Descripción

While exploring optimizations listed in the [documentation](https://huggingface.co/docs/diffusers/optimization/torch2.0), I find myself unable to free GPU memory after using `torch.compile` on a StableDiffusionXLPipeline UNet.

```Python
from diffusers import StableDiffusionXLPipeline

pipe = StableDiffusionXLPipeline.from_pretrained(
'stabilityai/stable-diffusion-xl-base-1.0',
torch_dtype=torch.float16,
variant="fp16",
use_safetensors=True
).to('cuda')

# Compile UNet
pipe.unet = torch.compile(pipe.unet, mode="reduce-overhead", fullgraph=True)

generator = torch.Generator(device="cuda").manual_seed(42)
prompt = "a photo of an astronaut riding a horse on mars"

image = pipe(prompt=prompt, num_inference_steps=20, generator=generator).images[0]

del pipe

gc.collect()
torch._dynamo.reset()
torch.cuda.empty_cache()
torch.cuda.synchronize()

# GPU memory is still in use, but it's not the case when we do not compile the pipeline unet.
```

It can sometimes be useful to free the GPU memory, especially if you want to load and compile another pipeline checkpoint to perform another large number of generations.

I made a [code reproduction](https://colab.research.google.com/drive/191XutMhBarF0MFXOmPbsOQH-a8CB9vmE?usp=drive_link) in collab for testing.

Am I missing something? Could it be a [memory leak](https://dev-discuss.pytorch.org/t/fixing-torch-compile-reference-leaks-automatic-deletion-of-dynamo-code-objects/2197) on the compilation backend side, in which case it might be better to turn to PyTorch to discuss about this?

### System Info
python: 3.10.12
diffusers: 0.30.3
torch: 2.4.1+cu121
Running on Google Colab?: Yes

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Start with the linked Colab reproduction and compare GPU memory cleanup with and without torch.compile on the StableDiffusionXLPipeline UNet. Check the referenced torch.compile reference-leak discussion and determine whether the behavior belongs in diffusers or PyTorch. Done means identifying the cause and recording a confirmed fix or workaround.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python, pytorch
Área
machine-learning, performance
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.