NVIDIA-NeMo / NVIDIA-NeMo/RL

On-policy distillation with different tokenizers

Open
#1,827 0 comments 0 reactions 1 assignee Claimed by @sharathts View on GitHub
enhancement x-shop
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

similar to https://huggingface.co/docs/trl/en/gold_trainer

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.