deepspeedai / deepspeedai/DeepSpeed

[BUG] DeepSpeedZeroOptimizer_Stage3: It cannot reduce the gradients remained in the bucket timely

Open
#4,879 2 comments 0 reactions 1 assignee View on GitHub

@GuanhuaWang is already working on this.

Since Jan 19, 2024.

bug training
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Describe the bug
After each step, when gradient_accumulation_steps is set to be 1 and in the end of each step, __reduce_and_partition_ipg_grads should be invoked to reduce the remainning gradients in the bucket. Right now, it does not. It can cause the following issue:
At the end of each step, only partial gradients are reduced, while the remained gradients are not reduced.

To Reproduce
You could refer to independent_gradient_partition_epilogue, which is not invoked.
Also, there is another bug that gradient_accumulation_steps is not used in stage3.py.
It is easy to reproduce, which you can use any model to find out the issue.

Expected behavior
At the beginning of each step, "self.elements_in_ipg_bucket" should be 0, while it is not right now.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.