Add agent_ref level metrics in training for GRPO and OPD
Open
enhancement
Feature
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
For current GRPO and OPD runs, the logprob gap / reward variacne / advantage etc seem to be reported and tracked only in aggregate instead of per-environment (from Nemo-Gym). It would be great to add agent_ref level metrics.
Contributor guide
Assessment
This issue has not been assessed yet.