THUDM / THUDM/slime

Fused kernel DPO loss

Open
#146 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

When we were training RL with Verl, we found that the fused-kernel DPO loss is very effective for long contexts. It would be great if Slime could support this. I'm trying to figure out how to implement it with Megatron, but it's quite complicated. I hope someone can contribute to this.
Of course, I'm still working on it, and I’ll be happy to share it if I can make it work. Thanks.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are named. Start by tracing Slime’s existing DPO and Megatron training paths, then compare them with the fused-kernel approach used in Verl. Done means Slime supports fused-kernel DPO loss for long-context RL training, with validation of the resulting training behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.