NVIDIA-NeMo / NVIDIA-NeMo/RL

Break up grpo_train to be more composable

Open
#706 0 comments 0 reactions 0 assignees View on GitHub
enhancement external x-futurehouse x-google
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

Break up grpo_train to be more functional/composable to enable R&D
Then implementing e.g. REINFORCE becomes a matter of swapping out loss function, rollout config
Also lets us easily experiment with ideas like expert demonstrations, curriculum learning, etc. These Composable primitives we can assemble into GRPO, REINFORCE, PPO, …

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.