NVIDIA-NeMo / NVIDIA-NeMo/RL

Add agent_ref level metrics in training for GRPO and OPD

Open
#2,537 0 comments 0 reactions 1 assignee Claimed by @terrykong View on GitHub
enhancement Feature
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

For current GRPO and OPD runs, the logprob gap / reward variacne / advantage etc seem to be reported and tracked only in aggregate instead of per-environment (from Nemo-Gym). It would be great to add agent_ref level metrics.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.