Fused kernel DPO loss
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.5k
- Forks
- 1.3k
- Avg merge
- 5h 36m
- Merged PRs (30d)
- 22
Description
When we were training RL with Verl, we found that the fused-kernel DPO loss is very effective for long contexts. It would be great if Slime could support this. I'm trying to figure out how to implement it with Megatron, but it's quite complicated. I hope someone can contribute to this.
Of course, I'm still working on it, and I’ll be happy to share it if I can make it work. Thanks.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are named. Start by tracing Slime’s existing DPO and Megatron training paths, then compare them with the fused-kernel approach used in Verl. Done means Slime supports fused-kernel DPO loss for long-context RL training, with validation of the resulting training behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100