deepseek-ai / deepseek-ai/DeepSeek-Math
iterative RL里训reward model的数据为啥能迭代优化?
Open
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 592
- PR merge metrics
- No merged PRs in 30d
Description
是人工参与进来标注?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.