huggingface / huggingface/diffusers

Slow SDXL inference with JAX on Cloud TPU v5e for sizes other than 1024x1024

オープン
#6,882 コメント 7 件 リアクション 1 件 担当者 0 名 GitHub で見る
bug jax/flax stale
主要言語
Python
スター
34.5k
フォーク
7.3k
平均マージ
3日 3時間
マージ済み PR(30日)
91

説明

### Describe the bug

Following the blog post on Accelerating Stable Diffusion XL Inference with JAX on Cloud TPU v5e. This worked magically until I tried to generate an image in a different size. At 1024x1024 we get inference latency of ~3s per image (as compared to ~8s on the NVIDIA A10G). But change the resolution to 1280x960 and we see next to no improvement.

### Reproduction

Use the same code as in the blog post: https://huggingface.co/blog/sdxl_jax

Changes:

```
def generate(
prompt,
negative_prompt,
seed=default_seed,
guidance_scale=default_guidance_scale,
num_inference_steps=default_num_steps,
width=1024,
height=1024,
):
prompt_ids, neg_prompt_ids = tokenize_prompt(prompt, negative_prompt)
prompt_ids, neg_prompt_ids, rng = replicate_all(prompt_ids, neg_prompt_ids, seed)
images = pipeline(
prompt_ids,
p_params,
rng,
num_inference_steps=num_inference_steps,
neg_prompt_ids=neg_prompt_ids,
guidance_scale=guidance_scale,
width=width,
height=height,
jit=True,
).images

# convert the images to PIL
images = images.reshape((images.shape[0] * images.shape[1], ) + images.shape[-3:])
return pipeline.numpy_to_pil(np.array(images))
```

```
start = time.time()
print(f"Compiling ...")
generate(default_prompt, default_neg_prompt, width=960, height=1280)
print(f"Compiled in {time.time() - start}")
```

```
start = time.time()
print("starting")
prompt = "llama in ancient Greece, oil on canvas"
neg_prompt = "cartoon, illustration, animation"
images = generate(prompt, neg_prompt, width=960, height=1280)
print(f"Inference in {time.time() - start}")
```

### Logs

_No response_

### System Info

Python: 3.10.6
Diffusers: 0.26.2
Torch: 2.2.0+cu121
Jax: 0.4.23
Flax: 0.8.0

### Who can help?

@patrickvonplaten @yiyixuxu @DN6

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

generate関数と、再現手順にリンクされているSDXL JAXのブログ記事のコードから始めてください。提供された環境を使用して、1024x1024と960x1280のコンパイル時間および推論時間を比較し、その後、正方形でない次元がどのように処理されるかを追跡してください。両方の解像度について測定を行い、TPUの速度向上が見られない原因を特定して修正できれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
machine-learning, performance
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。