huggingface / huggingface/diffusers
[Training] Resume checkpoint global step inconsistent/confusion across scripts
- 主要言語
- Python
- スター
- 34.5k
- フォーク
- 7.3k
- 平均マージ
- 3日 3時間
- マージ済み PR(30日)
- 91
説明
### Describe the bug
Hi,
I have been working on training scripts for multiple models (T2I, IP2P) and found the different logic to calculate `step` and `epoch` while resuming training different across scripts.
In `train_text_to_image.py` script [link](https://github.com/huggingface/diffusers/blob/fe5f035f797a5fa663a98030c9d0ec2f982cd09d/examples/text_to_image/train_text_to_image.py#L902)
```
accelerator.print(f"Resuming from checkpoint {path}")
accelerator.load_state(os.path.join(args.output_dir, path))
global_step = int(path.split("-")[1])
initial_global_step = global_step
first_epoch = global_step // num_update_steps_per_epoch
```
In `train_instruct_pix2pix.py` script [link](https://github.com/huggingface/diffusers/blob/fe5f035f797a5fa663a98030c9d0ec2f982cd09d/examples/instruct_pix2pix/train_instruct_pix2pix.py#L825)
```
accelerator.print(f"Resuming from checkpoint {path}")
accelerator.load_state(os.path.join(args.output_dir, path))
global_step = int(path.split("-")[1])
resume_global_step = global_step * args.gradient_accumulation_steps
first_epoch = global_step // num_update_steps_per_epoch
resume_step = resume_global_step % (num_update_steps_per_epoch * args.gradient_accumulation_steps)
```
In the similar [issue](https://github.com/huggingface/diffusers/issues/5005), some changes are made for the progress bar inconsistency but I am bit confused with the following things:-
1. The multiplication of `args.gradient_accumulation_steps` in `train_instruct_pix2pix.py` script
2. In general, when does global-step indicate and how does it's being updated, in both the scripts I can see the following code but couldn't understand it from `accelerate` documentation
```
if accelerator.sync_gradients:
if args.use_ema:
ema_unet.step(unet.parameters())
progress_bar.update(1)
global_step += 1
accelerator.log({"train_loss": train_loss}, step=global_step)
train_loss = 0.0
```
If we are using multiple GPUs with gradient accumulation, at what event `global_step` is updated- is it being updated independently by each GPU (since the code is not wrapped with `accelerator.is_main_process`), also how accumulation affecting the tracking here?
### Reproduction
-
### Logs
_No response_
### System Info
-
### Who can help?
@sayakpaul
コントリビューションガイド
調査の方向性
Start by comparing the resume logic in examples/text_to_image/train_text_to_image.py and examples/instruct_pix2pix/train_instruct_pix2pix.py, then read the accelerator.sync_gradients block that updates global_step. Trace how gradient accumulation and multiple processes affect these values, and define consistent resume and progress-tracking behavior across both scripts.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python, pytorch
- 領域
- machine-learning
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 28/100