huggingface / huggingface/diffusers

[Training] Resume checkpoint global step inconsistent/confusion across scripts

オープン
#8,296 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る
bug stale
主要言語
Python
スター
34.5k
フォーク
7.3k
平均マージ
3日 3時間
マージ済み PR(30日)
91

説明

### Describe the bug

Hi,
I have been working on training scripts for multiple models (T2I, IP2P) and found the different logic to calculate `step` and `epoch` while resuming training different across scripts.
In `train_text_to_image.py` script [link](https://github.com/huggingface/diffusers/blob/fe5f035f797a5fa663a98030c9d0ec2f982cd09d/examples/text_to_image/train_text_to_image.py#L902)
```
accelerator.print(f"Resuming from checkpoint {path}")
accelerator.load_state(os.path.join(args.output_dir, path))
global_step = int(path.split("-")[1])
initial_global_step = global_step
first_epoch = global_step // num_update_steps_per_epoch
```
In `train_instruct_pix2pix.py` script [link](https://github.com/huggingface/diffusers/blob/fe5f035f797a5fa663a98030c9d0ec2f982cd09d/examples/instruct_pix2pix/train_instruct_pix2pix.py#L825)
```
accelerator.print(f"Resuming from checkpoint {path}")
accelerator.load_state(os.path.join(args.output_dir, path))
global_step = int(path.split("-")[1])

resume_global_step = global_step * args.gradient_accumulation_steps
first_epoch = global_step // num_update_steps_per_epoch
resume_step = resume_global_step % (num_update_steps_per_epoch * args.gradient_accumulation_steps)
```
In the similar [issue](https://github.com/huggingface/diffusers/issues/5005), some changes are made for the progress bar inconsistency but I am bit confused with the following things:-
1. The multiplication of `args.gradient_accumulation_steps` in `train_instruct_pix2pix.py` script
2. In general, when does global-step indicate and how does it's being updated, in both the scripts I can see the following code but couldn't understand it from `accelerate` documentation
```
if accelerator.sync_gradients:
if args.use_ema:
ema_unet.step(unet.parameters())
progress_bar.update(1)
global_step += 1
accelerator.log({"train_loss": train_loss}, step=global_step)
train_loss = 0.0
```
If we are using multiple GPUs with gradient accumulation, at what event `global_step` is updated- is it being updated independently by each GPU (since the code is not wrapped with `accelerator.is_main_process`), also how accumulation affecting the tracking here?

### Reproduction

-

### Logs

_No response_

### System Info

-

### Who can help?

@sayakpaul

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

Start by comparing the resume logic in examples/text_to_image/train_text_to_image.py and examples/instruct_pix2pix/train_instruct_pix2pix.py, then read the accelerator.sync_gradients block that updates global_step. Trace how gradient accumulation and multiple processes affect these values, and define consistent resume and progress-tracking behavior across both scripts.

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python, pytorch
領域
machine-learning
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
28/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。