NVIDIA-NeMo / NVIDIA-NeMo/RL

Scale RL features

Open
#1,475 0 comments 0 reactions 0 assignees View on GitHub
algorithm enhancement
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

- [ ] advantage normalization (batch level)
- [ ] zero-variance filtering
- [ ] FP32 lm head (train and gen)
- [ ] CISPO
- [ ] no-positive resampling (data curriculum: skip >=0.9 pass rate)

cc/ @parthchadha

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.