NVIDIA-NeMo / NVIDIA-NeMo/RL

[Feature Request] Add Closed-Loop Critic Refinement and Alignment

Open
#828 0 comments 0 reactions 0 assignees View on GitHub
algorithm external x-abeja
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

Kimi K2 paper proposed _Closed-Loop Critic Refinement and Alignment_ (During RL training, the critic model is refined using verifiable signals).

[https://arxiv.org/html/2507.20534v1#:~:text=Closed-Loop Critic Refinement and Alignment](https://arxiv.org/html/2507.20534v1#:~:text=Closed%2DLoop%20Critic%20Refinement%20and%20Alignment)

The paper argued
"Consequently, this holistic alignment yields comprehensive performance improvements across a wide spectrum of domains, including user intent understanding, creative writing, complex reasoning, and nuanced language comprehension."

Are there any planned add methods like these in the future?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.