huggingface / huggingface/diffusers
pipeline does not work with conversion to fp16 in the to cuda call
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 34.5k
- Forks
- 7.3k
- Avg merge
- 3d 3h
- Merged PRs (30d)
- 91
Description
Hi,
i was using and doing some fine-tuning with the Pixart model.
and i found a problem that i could pin it down to:
this code works:
pipe2 = PixArtAlphaPipeline.from_pretrained("PixArt-alpha/PixArt-XL-2-512x512", torch_dtype=torch.float16)
pipe2.to("cuda")
image = pipe2(prompt, num_inference_steps=20).images[0]
while this does not, it gives out just noise images :
pipe1 = PixArtAlphaPipeline.from_pretrained("PixArt-alpha/PixArt-XL-2-512x512")
pipe1.to("cuda", dtype=torch.float16)
image = pipe1(prompt, num_inference_steps=20).images[0]
i have looked through the code, and i cannot find any issues so far.
has anyone encountered a similar problem or has tips where to look further?
i am on WSL Ubuntu, diffusers==0.27.2, torch==2.2.2
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the two PixArtAlphaPipeline snippets with diffusers 0.27.2 and torch 2.2.2 on WSL Ubuntu, comparing the pipeline state and generated outputs after each CUDA conversion. Done means identifying why the two conversion paths differ and making the dtype conversion produce valid, non-noise images.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100