alibaba / alibaba/ROLL

ipro的训练代码中奖励函数是不可微的,训练中梯度不会更新,和论文中描述不一致,请问是哪里有问题呢?

Open
#300 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
312
Avg merge
1h 2m
Merged PRs (30d)
2

Description

辛苦解答一下

Contributor guide

No contributing guide indexed for this repository

Research direction

Locate the iPRO training code and trace the reward calculation through the loss and optimizer update. Compare the implementation with the paper's description, then reproduce the reported missing-gradient behavior. Done means identifying the cause of the discrepancy and confirming the expected gradient behavior or documenting why the reward is intentionally non-differentiable.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.