deepspeedai / deepspeedai/DeepSpeed
[BUG] Does deepspeed use early recomputation strategy?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Describe the bug
I saw that Gpipe used to recalculate the activation value of the next stage in advance to reduce delays, but when I used deepspeed training myself, I drew a timeline diagram of each stage. The backward time is about 3 times the forward time and I found that the early recomputation strategy is not used.
In the figure, purple represents backward and green represents forward. The depth of the pipeline is 3, and the gradient accumulation is 4.
The picture below is the early recomputation strategy of Gpipe that I saw in the paper Merak.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file, test, or entry point is named. Start by reproducing the reported three-stage, four-gradient-accumulation timeline and tracing the pipeline scheduling path that controls forward, backward, and recomputation ordering. Done requires confirming whether early recomputation is supported and documenting or correcting the behavior based on that finding.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100