NVIDIA / NVIDIA/Megatron-LM

Feature Request: adopt ant's AMem as an optional API to offload NCCL memory in RL scenario

Open
#2,423 1 comment 2 reactions 0 assignees View on GitHub
enhancement module: rl
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 3h
Merged PRs (30d)
272

Description

**Is your feature request related to a problem? Please describe.**
A clear and concise description of what the problem is. Ex. I'm always frustrated when

https://github.com/inclusionAI/asystem-amem Amen is an open source NCCL plugin for fast offload/reload, which could potential save up to 10GB memory saving for rollout.

**Describe the solution you'd like**
Add an optional dependency with Amem to build communication groups.

Contributor guide

Open the contributing guide

Research direction

The issue does not name files or tests. Start by locating how Megatron-LM builds communication groups and declares optional dependencies, then determine where an AMem-backed path would fit for RL rollout memory offload. Done means the optional dependency can be enabled without affecting existing users and the communication groups use it when configured.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.