alibaba / alibaba/UI-Ins

About Instruction-as-Reasoning Ablation Experiment

Open
#11 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
81
Forks
6
PR merge metrics
No merged PRs in 30d

Description

Dear authors,

Thank you for your insightfull work! I have some questions regarding the ablation experiments in the paper. **Table 4** presents two ablation experiments, distinguished by whether inference was performed during training. However, why do all benchmark scores in the last row of both experiments are same?

Does that mean, for ` SFT+RL `, there's no difference whether we reasoning in training stages or not?

Any response would be appreciated!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading Table 4 and comparing the two ablation setups described in the issue, especially the SFT+RL row and whether inference occurs during training. Done means providing an author-confirmed explanation for why the benchmark scores match and whether reasoning during training changes SFT+RL.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.