DPO dataset format and loss
Open
- Dominant language
- Python
- Stars
- 5.2k
- Forks
- 448
- Avg merge
- 3d 15h
- Merged PRs (30d)
- 26
Description
Should be quite easy to add for someone who knows the codebase. The biggest problem might be a new dataset format.
Don't expect I need to link this but it's pretty nice implementation of the loss:
https://github.com/huggingface/trl/blob/main/trl/trainer/dpo_trainer.py#L817
Contributor guide
Assessment
This issue has not been assessed yet.