关于dpo训练时chat template的使用问题
Open
- Dominant language
- Python
- Stars
- 5.2k
- Forks
- 448
- Avg merge
- 3d 15h
- Merged PRs (30d)
- 26
Description
你好,看了最新的提交,支持了dpo训练。但是从代码来看,似乎在对偏好数据集的处理时,并没有使用模型对应的chat template。从搜索到的资料来看,似乎使用与不使用的情况都存在。想请问下有试验过在chat模型上两种方式的差异吗?
Contributor guide
Research direction
Start by reading the DPO preference-dataset processing code introduced in the latest commits and trace whether the model's chat template is applied for chat models. Compare the behavior with and without the template, then document the expected approach and add coverage if the project identifies a preferred behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 18/100