ipro的训练代码中奖励函数是不可微的,训练中梯度不会更新,和论文中描述不一致,请问是哪里有问题呢?
Open
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 312
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 2
Description
辛苦解答一下
Contributor guide
No contributing guide indexed for this repository
Research direction
Locate the iPRO training code and trace the reward calculation through the loss and optimizer update. Compare the implementation with the paper's description, then reproduce the reported missing-gradient behavior. Done means identifying the cause of the discrepancy and confirming the expected gradient behavior or documenting why the reward is intentionally non-differentiable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100