deepseek-ai / deepseek-ai/DeepSeek-Math
关于强化学习GRPO训练的问题
Open
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 592
- PR merge metrics
- No merged PRs in 30d
Description

文中给出了强化学习部分的伪代码,但从里面看不出这里的一个step是一轮对话里面的一次问答,还是一次问答中输出的前后两个token。由于源码没有公开,虽然倾向于前者但仍不能肯定。烦请解答,谢谢!
Contributor guide
No contributing guide indexed for this repository
Research direction
No source file, test, or entry point is named; start with the GRPO pseudocode and image included in the issue. Clarify whether one step means a dialogue-level question-and-answer or adjacent output tokens, and consider the issue done when a maintainer provides an authoritative explanation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100