bethgelab / bethgelab/sober-reasoning
Separate SFT+RL from RL only
- Dominant language
- HTML
- Stars
- 92
- Forks
- 6
- PR merge metrics
- No merged PRs in 30d
Description
Hi!
I think that models trained with distillation followed by reinforcement learning, or multiple distillation steps, should have a separate section (or at least an “SFT + RL” indication).
For instance for Qwen2.5-Math-7B, the Sky-T1-7B model is trained with 4-step SFT->RL->SFT->RL vs Oat-Zero and others which are trained in a single RL run.
Wdyt?
(aside: I think the leaderboard would benefit greatly from a verified / unverified section like SWE-bench so that new releases and be added and compared quickly. It would need an easy way to run the full pipeline locally, but I think this would be very useful to the community.)
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, or entry points. Start by locating the leaderboard data and model-training metadata, then clarify whether the intended scope is only an SFT + RL indication or also the proposed verified/unverified workflow; done means the distinction and its presentation are agreed and consistently represented.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100