huggingface / huggingface/diffusers
Freeing GPU memory after `torch.compile` StableDiffusionXLPipeline UNet
- Lingua principale
- Python
- Stelle
- 34.5k
- Fork
- 7.3k
- Merge medio
- 3g 3h
- PR unite (30g)
- 91
Descrizione
While exploring optimizations listed in the [documentation](https://huggingface.co/docs/diffusers/optimization/torch2.0), I find myself unable to free GPU memory after using `torch.compile` on a StableDiffusionXLPipeline UNet.
```Python
from diffusers import StableDiffusionXLPipeline
pipe = StableDiffusionXLPipeline.from_pretrained(
'stabilityai/stable-diffusion-xl-base-1.0',
torch_dtype=torch.float16,
variant="fp16",
use_safetensors=True
).to('cuda')
# Compile UNet
pipe.unet = torch.compile(pipe.unet, mode="reduce-overhead", fullgraph=True)
generator = torch.Generator(device="cuda").manual_seed(42)
prompt = "a photo of an astronaut riding a horse on mars"
image = pipe(prompt=prompt, num_inference_steps=20, generator=generator).images[0]
del pipe
gc.collect()
torch._dynamo.reset()
torch.cuda.empty_cache()
torch.cuda.synchronize()
# GPU memory is still in use, but it's not the case when we do not compile the pipeline unet.
```
It can sometimes be useful to free the GPU memory, especially if you want to load and compile another pipeline checkpoint to perform another large number of generations.
I made a [code reproduction](https://colab.research.google.com/drive/191XutMhBarF0MFXOmPbsOQH-a8CB9vmE?usp=drive_link) in collab for testing.
Am I missing something? Could it be a [memory leak](https://dev-discuss.pytorch.org/t/fixing-torch-compile-reference-leaks-automatic-deletion-of-dynamo-code-objects/2197) on the compilation backend side, in which case it might be better to turn to PyTorch to discuss about this?
### System Info
python: 3.10.12
diffusers: 0.30.3
torch: 2.4.1+cu121
Running on Google Colab?: Yes
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Start with the linked Colab reproduction and compare GPU memory cleanup with and without torch.compile on the StableDiffusionXLPipeline UNet. Check the referenced torch.compile reference-leak discussion and determine whether the behavior belongs in diffusers or PyTorch. Done means identifying the cause and recording a confirmed fix or workaround.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python, pytorch
- Ambito
- machine-learning, performance
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 35/100