NVIDIA-NeMo / NVIDIA-NeMo/RL

Colocated training state offload is inefficient

Open
#3,976 3 comments 0 reactions 0 assignees View on GitHub
bug Speed
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

**Describe the bug**

The optional offload of training state during colocated training is too slow to be useful.

This can be optimized upstream in NVIDIA/Megatron-LM, such as with NVIDIA/Megatron-LM#4241. A proper triage of the problem still needs to be made.

**Steps/Code to reproduce bug**

Please list *minimal* steps or code snippet for us to be able to reproduce the bug.

A helpful guide on on how to craft a minimal bug report http://matthewrocklin.com/blog/work/2018/02/28/minimal-bug-reports.

**Expected behavior**

A clear and concise description of what you expected to happen.

**Additional context**

Add any other context about the problem here.

Contributor guide

Open the contributing guide

Research direction

The issue names no project file, test, or entry point. Start by locating the colocated training state offload implementation and comparing the reported inefficiency with NVIDIA/Megatron-LM#4241, then establish a minimal reproduction. Done means the slowdown is properly triaged with an identified bottleneck and a clear next step.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.