bethgelab / bethgelab/sober-reasoning

Separate SFT+RL from RL only

Open
#17 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
HTML
Stars
92
Forks
6
PR merge metrics
No merged PRs in 30d

Description

Hi!

I think that models trained with distillation followed by reinforcement learning, or multiple distillation steps, should have a separate section (or at least an “SFT + RL” indication).

For instance for Qwen2.5-Math-7B, the Sky-T1-7B model is trained with 4-step SFT->RL->SFT->RL vs Oat-Zero and others which are trained in a single RL run.

Wdyt?

(aside: I think the leaderboard would benefit greatly from a verified / unverified section like SWE-bench so that new releases and be added and compared quickly. It would need an easy way to run the full pipeline locally, but I think this would be very useful to the community.)

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Start by locating the leaderboard data and model-training metadata, then clarify whether the intended scope is only an SFT + RL indication or also the proposed verified/unverified workflow; done means the distinction and its presentation are agreed and consistently represented.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.