NVIDIA-NeMo / NVIDIA-NeMo/RL

Questions about nemotron training

Open
#2,944 0 comments 0 reactions 0 assignees View on GitHub
community-request enhancement waiting-on-maintainers
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

Hi, thank you for releasing the high quality code and ckpts. I have some questions for reproducing nemotron ultra 3 training. As the nemotron repo is no longer active, I propose questions in this repo.

1. For the model use teacher, the information is very vague. It is better to know the details of model training, e.g., sft or rl, and using what kinds of datasets.

Image

2. The section mentions the usage of "general sft data", but I cannot find the description in the manuscript:

Image

Is this same as the general sft data?

Thanks, having these details will be really helpful!

Contributor guide

Open the contributing guide

Research direction

Start with the Nemotron training section referenced in the issue and compare its “general sft data” wording with the manuscript and the attached figures. Done means documenting the teacher model’s SFT or RL method and datasets, and clarifying whether “general sft data” is the same data described in the manuscript.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.