huggingface / huggingface/diffusers

`prepare_ip_adapter_image_embeds` Bug Causes Feature Mixing During Batch Processing in IP-Adapter

Open
#9,813 4 comments 1 reaction 2 assignees Claimed by @asomoza View on GitHub
bug
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

### Describe the bug

**The `prepare_ip_adapter_image_embeds ` function has a bug that results in unintended feature mixing across images during batch processing. This issue causes the generated images to combine features from multiple reference images, instead of maintaining a one-to-one correspondence with each reference.**

When using the pipeline in batch mode, I use `ip_adapter_image_embeds` with a shape of `(2*B, N, C)` and set `num_images_per_prompt=1`. I expect the pipeline to generate `B` images, where each generated image should correspond directly to each reference in `ip_adapter_image_embeds` (note that `2*B` includes the negative image embedding for classifier-free guidance).

https://github.com/huggingface/diffusers/blob/31058cdaef63ca660a1a045281d156239fba8192/src/diffusers/pipelines/stable_diffusion/pipeline_stable_diffusion.py#L950-L957

https://github.com/huggingface/diffusers/blob/9a92b8177cb3f8bf4b095fff55da3b45a3607960/src/diffusers/pipelines/stable_diffusion/pipeline_stable_diffusion.py#L561-L569

However, when processing `ip_adapter_image_embeds` in the pipeline, the tensor gets duplicated `num_images_per_prompt * batch_size = 1 * B` times. This leads to the `image_embeds ` tensor having a shape of `(B*2*B, N, C)` instead of the expected shape of `(2*B, N, C)` .

In the `IPAdapterAttnProcessor2_0` class, the view operation is applied to the input image_embeds tensor. This prevents a shape mismatch error, but it leads to `ip_key` and `ip_value` containing mixed features from multiple reference images. As a result, the features of the generated images are a mixture of several reference images instead of having a one-to-one correspondence.

https://github.com/huggingface/diffusers/blob/9a92b8177cb3f8bf4b095fff55da3b45a3607960/src/diffusers/models/attention_processor.py#L4112-L4122

**Although I temporarily resolved the issue by changing the `num_images_per_prompt*batch_size` parameter passed to the `prepare_ip_adapter_image_embeds` method to `num_images_per_prompt`, could this potentially cause issues in other scenarios?**
```python
if ip_adapter_image is not None or ip_adapter_image_embeds is not None:
image_embeds = self.prepare_ip_adapter_image_embeds(
ip_adapter_image,
ip_adapter_image_embeds,
device,
num_images_per_prompt,
self.do_classifier_free_guidance,
)
```

### Reproduction

Here’s a demo script that illustrates the issue. The script loads two reference images (image1 and image2), extracts their embeddings, and uses them as input to the pipeline in batch mode.
```python
import torch
from diffusers import StableDiffusionPipeline, DDIMScheduler
from diffusers.utils import load_image
from insightface.app import FaceAnalysis
import cv2
import numpy as np

pipeline = StableDiffusionPipeline.from_pretrained(
"../checkpoints/Realistic_Vision_V4.0_noVAE", # Replace with your model weights path
torch_dtype=torch.float16,
).to("cuda")

pipeline.scheduler = DDIMScheduler.from_config(pipeline.scheduler.config)
pipeline.load_ip_adapter("../checkpoints/IP-Adapter", subfolder=None,
weight_name="ip-adapter-faceid_sd15.bin", image_encoder_folder=None) #Replace with your model weights path
pipeline.set_ip_adapter_scale(1.0)

#Replace with your model weights path
app = FaceAnalysis(name="/root/data1/IP-Face/checkpoints/insightface", providers=['CUDAExecutionProvider', 'CPUExecutionProvider']) #Replace with your model weights path
app.prepare(ctx_id=0, det_size=(384, 384))

image1 = load_image('../test_image/65.jpg')
image2 = load_image('../test_image/27022.jpg')

face1 = cv2.cvtColor(np.asarray(image1),cv2.COLOR_RGB2BGR)
face1 = app.get(face1)
face1_embedding = torch.from_numpy(face1[0].normed_embedding)
face1_embedding = face1_embedding.reshape(1,1,-1)

face2 = cv2.cvtColor(np.asarray(image2),cv2.COLOR_RGB2BGR)
face2 = app.get(face2)
face2_embedding = torch.from_numpy(face2[0].normed_embedding)
face2_embedding = face2_embedding.reshape(1,1,-1)

ref_face_embedding = torch.cat([face1_embedding,face2_embedding])
neg_ref_face_embedding = torch.zeros_like(ref_face_embedding)

batch_id_embeds = torch.cat([neg_ref_face_embedding, ref_face_embedding]).to(dtype=torch.float16, device="cuda")
batch_size = int(batch_id_embeds.shape[0]/2)

generator = torch.Generator(device="cpu").manual_seed(2023)
images = pipeline(
prompt=["photo of a woman in red dress in a garden"]*batch_size,
ip_adapter_image_embeds=[batch_id_embeds],
negative_prompt=["monochrome, lowres, bad anatomy, worst quality, low quality"]*batch_size,
num_inference_steps=50, num_images_per_prompt=1,
generator=generator
).images

```
## Reference Images
The reference images image1 and image2 used as input embeddings:
| image1| image2|
| --- | --- |
| ![image1](https://github.com/user-attachments/assets/cad95244-1503-4185-abd7-b42265cac40f) | ![image2](https://github.com/user-attachments/assets/f368eab5-fbad-45b9-9f8f-a43d20519423) |

## Generated Images in Batch Mode
Using the demo code above, the following images were generated. These images exhibit features mixed from both references instead of corresponding uniquely to one.
| Generated Image 1| Generated Image 2|
| --- | --- |
|![issue_1](https://github.com/user-attachments/assets/1e8f262b-fe9b-4104-a5b2-9383a8fcd36b) | ![issue_2](https://github.com/user-attachments/assets/3f0241c9-d1bb-4aac-8828-9ea6a31513dd) |

## Expected Behavior
In single-image processing (non-batch mode), the pipeline works as expected, producing distinct images for each reference:
| Generated Image 1(no Batch)| Generated Image 2(no Batch)|
| --- | --- |
|![issue_3](https://github.com/user-attachments/assets/63d052a5-eadd-48c3-875e-077fe15bd32f) | ![issue_4](https://github.com/user-attachments/assets/a0570702-72b1-473f-b6de-ce9cd8349e33) |

### Logs

_No response_

### System Info

- diffusers == 0.30.3
- torch == 2.4.1+cu121
- insightface == 0.7.3
- python == 3.10.0

### Who can help?

@asomoza

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.