Colocated training state offload is inefficient
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
**Describe the bug**
The optional offload of training state during colocated training is too slow to be useful.
This can be optimized upstream in NVIDIA/Megatron-LM, such as with NVIDIA/Megatron-LM#4241. A proper triage of the problem still needs to be made.
**Steps/Code to reproduce bug**
Please list *minimal* steps or code snippet for us to be able to reproduce the bug.
A helpful guide on on how to craft a minimal bug report http://matthewrocklin.com/blog/work/2018/02/28/minimal-bug-reports.
**Expected behavior**
A clear and concise description of what you expected to happen.
**Additional context**
Add any other context about the problem here.
Contributor guide
Research direction
The issue names no project file, test, or entry point. Start by locating the colocated training state offload implementation and comparing the reported inefficiency with NVIDIA/Megatron-LM#4241, then establish a minimal reproduction. Done means the slowdown is properly triaged with an identified bottleneck and a clear next step.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100