google-deepmind / google-deepmind/dm_control

Reward is always 0 in Hopper

Abierto
#411 1 comentario 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
4.7k
Forks
764
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

In Hopper, reward is always 0 independent of the action. The problem is that this code which is used in reward calculation in hopper.py get_reward(self, physics) function, always return 0 as the value of standing:
`standing = rewards.tolerance(physics.height(), (_STAND_HEIGHT, 2))`

When I check `env._physics.height()` after environment reset, I saw that the height is initialized as a negative value always. Therefore it is impossible for this code fragment `rewards.tolerance(physics.height(), (_STAND_HEIGHT, 2))` to return 1, since _STAND_HEIGHT is defined as 0.6 for Hopper.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.