InternLM / InternLM/xtuner

DPO dataset format and loss

Open
#344 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.2k
Forks
448
Avg merge
3d 15h
Merged PRs (30d)
26

Description

Should be quite easy to add for someone who knows the codebase. The biggest problem might be a new dataset format.

Don't expect I need to link this but it's pretty nice implementation of the loss:
https://github.com/huggingface/trl/blob/main/trl/trainer/dpo_trainer.py#L817

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.