NVIDIA-NeMo / NVIDIA-NeMo/RL

[super-pr] terminal pivot env in gym + reasoning on/off in gym + custom reasoning parser + skip calculating prev_logprob when force_onpolicy_ratio is true

Open
#1,916 0 comments 0 reactions 1 assignee Claimed by @HeyyyyyyG View on GitHub
super-v3
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

e476235662b9537f32bce78da08917df1c64158e

This squash commit can exclude the replay buffer changes since that'll happen as part of #1906

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.