google-deepmind / google-deepmind/dm_control
Reward is always 0 in Hopper
- 主要语言
- Python
- 星标
- 4.7k
- 派生
- 764
- PR 合并指标
- 30 天内没有已合并 PR
描述
In Hopper, reward is always 0 independent of the action. The problem is that this code which is used in reward calculation in hopper.py get_reward(self, physics) function, always return 0 as the value of standing:
`standing = rewards.tolerance(physics.height(), (_STAND_HEIGHT, 2))`
When I check `env._physics.height()` after environment reset, I saw that the height is initialized as a negative value always. Therefore it is impossible for this code fragment `rewards.tolerance(physics.height(), (_STAND_HEIGHT, 2))` to return 1, since _STAND_HEIGHT is defined as 0.6 for Hopper.
贡献指南
调研方向
从 hopper.py 中的 get_reward(self, physics) 开始,检查重置高度以及 rewards.tolerance 和 _STAND_HEIGHT。复现报告的重置后奖励为零的问题,确定站立状态为何仍为零,并验证修复后奖励会对动作作出适当响应。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- machine-learning
- Issue 类型
- 缺陷
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 38/100