deepspeedai / deepspeedai/DeepSpeed
Question: how to continue the training with more or fewer GPUs
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
If there are N GPUs, the snapshot will be N files for optimizer states. Each file corresponds to 1 GPU. (let me know if the understanding is not correct). Then, how to continue the training with more GPU, say, 2N GPUs? Is there an easy way to consolidate the optimizer states?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing DeepSpeed's snapshot and optimizer-state handling to verify the one-file-per-GPU behavior described in the issue. Done would require a clear, tested way to consolidate optimizer states and resume training when moving from N GPUs to 2N GPUs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100