NVIDIA-NeMo / NVIDIA-NeMo/RL

NeMo RL Software Architecture Update: PR Tracking

Open
#2,905 0 comments 0 reactions 0 assignees View on GitHub
Documentation
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

# NeMo RL Evolution: PR Tracking

This issue consolidates the open and in-flight PRs across the NeMo RL evolution workstreams. It is meant as a single index so we can track status in one place. PRs are grouped by workstream, with cross-workstream PRs noted where relevant.

## Data Plane

Owner: Zhiyu Li

- [x] NVIDIA-NeMo/RL#2439: Data Plane initial integration
- [ ] NVIDIA-NeMo/RL#2593: Async RL example (Data Plane with Simple Storage)
- [x] NVIDIA-NeMo/RL#2616: Nightly tests coverage (Data Plane with Simple Storage)
- [x] TransferQueue/TransferQueue#167: RayRDT with NIXL in TransferQueue

### Router Replay (R3)

Owner: Zeyu Zhou

- [x] NVIDIA-NeMo/RL#2590: R3 PR with docs and wandb report
- [ ] NVIDIA-NeMo/RL#2465: R3 base PR

### FT, Distillation

Owner: Pranav Prashant Thombre

- [ ] NVIDIA-NeMo/RL#2580: FT / Distillation PR
- [ ] NVIDIA-NeMo/RL#2492: Fault tolerance support in Data Plane for the Async controller

## Refit

Owner: Songlin Jiang

Core refactor series (blocks the delta and RDMA refit work):

- [x] NVIDIA-NeMo/RL#2466: WeightSynchronizer ABC with IPC/HTTP/NCCL transports (merged)
- [ ] NVIDIA-NeMo/RL#2467: wire synchronizer into GRPO and distillation
- [ ] NVIDIA-NeMo/RL#2470: unify policy phase transitions (PolicyTrainerInterface)
- [ ] NVIDIA-NeMo/RL#2475: generation backend registry and create_generation() factory (also Generation)

### Delta Weight Transfer

- [ ] NVIDIA-NeMo/RL#2444: delta weight transfer (perf optimization and data-race fix)

### P2P RDMA based refit

- [ ] NVIDIA-NeMo/RL#2608: checkpoint engine interface with NIXL backend PoC

### NCCL hierarchical API

Owner: Youngeun Kwon

- [ ] NVIDIA-NeMo/RL#2413: NCCL hierarchical API / RDMA based refit

## AsyncRL

Owners: Akash Mehra, Yuki Huang

Train pump (3-PR async-GRPO split-API stack):

- [ ] NVIDIA-NeMo/RL#2700: SingleController streaming train_pump (split-API consumer)
- [ ] NVIDIA-NeMo/RL#2692: DTensor v1/v2 split-API and PolicyTrainerActor
- [ ] NVIDIA-NeMo/RL#2683: Megatron split-API train-step state machine

Rollout and per-prompt streaming:

- [ ] NVIDIA-NeMo/RL#2566: native async rollout path (also Generation)
- [ ] NVIDIA-NeMo/RL#2567: unify rollout paths into a unified factory (also Generation)
- [ ] NVIDIA-NeMo/RL#2528: gym rollout path (also Gym and Generation)
- [ ] NVIDIA-NeMo/RL#2448: refactor async utils
- [ ] NVIDIA-NeMo/RL#2458: staleness-window support
- [ ] ray-project/enhancements#64: SingleController/HeadNode fault tolerance RFC
- [ ] ray-project/enhancements#65: SingleController/HeadNode fault tolerance RFC

Single Controller:
- [ ] NVIDIA-NeMo/RL#2819: support SingleController e2e run

## Gym

Owners: Ananth Subramaniam, Hemil Desai

- [ ] NVIDIA-NeMo/Gym#1368: Sandbox API large PR (to be split into smaller parts)
- [ ] NVIDIA-NeMo/RL#2528: gym rollout path (shared with AsyncRL)

## Generation

- [ ] NVIDIA-NeMo/RL #2930: Skip training for generation benchmarking
## Reference links (full URLs)

Contributor guide

Open the contributing guide

Research direction

Start by reading the checklist in this issue and opening the referenced NeMo RL, Gym, TransferQueue, and Ray enhancement PRs. Verify each workstream's status and ownership against those linked items. Done means the index accurately reflects which PRs are open or merged, with the listed workstreams and cross-workstream references kept current.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.