NeMo RL Software Architecture Update: PR Tracking
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
# NeMo RL Evolution: PR Tracking
This issue consolidates the open and in-flight PRs across the NeMo RL evolution workstreams. It is meant as a single index so we can track status in one place. PRs are grouped by workstream, with cross-workstream PRs noted where relevant.
## Data Plane
Owner: Zhiyu Li
- [x] NVIDIA-NeMo/RL#2439: Data Plane initial integration
- [ ] NVIDIA-NeMo/RL#2593: Async RL example (Data Plane with Simple Storage)
- [x] NVIDIA-NeMo/RL#2616: Nightly tests coverage (Data Plane with Simple Storage)
- [x] TransferQueue/TransferQueue#167: RayRDT with NIXL in TransferQueue
### Router Replay (R3)
Owner: Zeyu Zhou
- [x] NVIDIA-NeMo/RL#2590: R3 PR with docs and wandb report
- [ ] NVIDIA-NeMo/RL#2465: R3 base PR
### FT, Distillation
Owner: Pranav Prashant Thombre
- [ ] NVIDIA-NeMo/RL#2580: FT / Distillation PR
- [ ] NVIDIA-NeMo/RL#2492: Fault tolerance support in Data Plane for the Async controller
## Refit
Owner: Songlin Jiang
Core refactor series (blocks the delta and RDMA refit work):
- [x] NVIDIA-NeMo/RL#2466: WeightSynchronizer ABC with IPC/HTTP/NCCL transports (merged)
- [ ] NVIDIA-NeMo/RL#2467: wire synchronizer into GRPO and distillation
- [ ] NVIDIA-NeMo/RL#2470: unify policy phase transitions (PolicyTrainerInterface)
- [ ] NVIDIA-NeMo/RL#2475: generation backend registry and create_generation() factory (also Generation)
### Delta Weight Transfer
- [ ] NVIDIA-NeMo/RL#2444: delta weight transfer (perf optimization and data-race fix)
### P2P RDMA based refit
- [ ] NVIDIA-NeMo/RL#2608: checkpoint engine interface with NIXL backend PoC
### NCCL hierarchical API
Owner: Youngeun Kwon
- [ ] NVIDIA-NeMo/RL#2413: NCCL hierarchical API / RDMA based refit
## AsyncRL
Owners: Akash Mehra, Yuki Huang
Train pump (3-PR async-GRPO split-API stack):
- [ ] NVIDIA-NeMo/RL#2700: SingleController streaming train_pump (split-API consumer)
- [ ] NVIDIA-NeMo/RL#2692: DTensor v1/v2 split-API and PolicyTrainerActor
- [ ] NVIDIA-NeMo/RL#2683: Megatron split-API train-step state machine
Rollout and per-prompt streaming:
- [ ] NVIDIA-NeMo/RL#2566: native async rollout path (also Generation)
- [ ] NVIDIA-NeMo/RL#2567: unify rollout paths into a unified factory (also Generation)
- [ ] NVIDIA-NeMo/RL#2528: gym rollout path (also Gym and Generation)
- [ ] NVIDIA-NeMo/RL#2448: refactor async utils
- [ ] NVIDIA-NeMo/RL#2458: staleness-window support
- [ ] ray-project/enhancements#64: SingleController/HeadNode fault tolerance RFC
- [ ] ray-project/enhancements#65: SingleController/HeadNode fault tolerance RFC
Single Controller:
- [ ] NVIDIA-NeMo/RL#2819: support SingleController e2e run
## Gym
Owners: Ananth Subramaniam, Hemil Desai
- [ ] NVIDIA-NeMo/Gym#1368: Sandbox API large PR (to be split into smaller parts)
- [ ] NVIDIA-NeMo/RL#2528: gym rollout path (shared with AsyncRL)
## Generation
- [ ] NVIDIA-NeMo/RL #2930: Skip training for generation benchmarking
## Reference links (full URLs)
Contributor guide
Research direction
Start by reading the checklist in this issue and opening the referenced NeMo RL, Gym, TransferQueue, and Ray enhancement PRs. Verify each workstream's status and ownership against those linked items. Done means the index accurately reflects which PRs are open or merged, with the listed workstreams and cross-workstream references kept current.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100