[Alignment training (DPO)] Megatron support for QWEN2.5 series models
Open
- Dominant language
- Python
- Stars
- 450
- Forks
- 40
- PR merge metrics
- No merged PRs in 30d
Description
when will the QWEN2.5 series models support alignment training (RLHF、DPO、OnlineDPO、GRPO) using the Megatron
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing ChatLearn's existing alignment-training support and the entry points for Megatron and QWEN model integration. Determine the scope required for QWEN2.5 support across RLHF, DPO, OnlineDPO, and GRPO; done means these workflows operate with the QWEN2.5 series through Megatron.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100