NVIDIA-NeMo / NVIDIA-NeMo/RL

Support additional preference-based objectives

Open
#193 3 comments 0 reactions 1 assignee Claimed by @ashors1 View on GitHub
external x-tii
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

Aligner supports IPO and RPO in addition to DPO. We should support these objectives as well.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.