About Instruction-as-Reasoning Ablation Experiment
- Dominant language
- Python
- Stars
- 81
- Forks
- 6
- PR merge metrics
- No merged PRs in 30d
Description
Dear authors,
Thank you for your insightfull work! I have some questions regarding the ablation experiments in the paper. **Table 4** presents two ablation experiments, distinguished by whether inference was performed during training. However, why do all benchmark scores in the last row of both experiments are same?
Does that mean, for ` SFT+RL `, there's no difference whether we reasoning in training stages or not?
Any response would be appreciated!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading Table 4 and comparing the two ablation setups described in the issue, especially the SFT+RL row and whether inference occurs during training. Done means providing an author-confirmed explanation for why the benchmark scores match and whether reasoning during training changes SFT+RL.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100