InternLM / InternLM/xtuner

关于dpo训练时chat template的使用问题

Open
#857 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.2k
Forks
448
Avg merge
3d 15h
Merged PRs (30d)
26

Description

你好,看了最新的提交,支持了dpo训练。但是从代码来看,似乎在对偏好数据集的处理时,并没有使用模型对应的chat template。从搜索到的资料来看,似乎使用与不使用的情况都存在。想请问下有试验过在chat模型上两种方式的差异吗?

Contributor guide

Open the contributing guide

Research direction

Start by reading the DPO preference-dataset processing code introduced in the latest commits and trace whether the model's chat template is applied for chat models. Compare the behavior with and without the template, then document the expected approach and add coverage if the project identifies a preferred behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
18/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.