google-deepmind / google-deepmind/dm_control

Reward is always 0 in Hopper

Offen
#411 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
4.7k
Forks
764
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

In Hopper, reward is always 0 independent of the action. The problem is that this code which is used in reward calculation in hopper.py get_reward(self, physics) function, always return 0 as the value of standing:
`standing = rewards.tolerance(physics.height(), (_STAND_HEIGHT, 2))`

When I check `env._physics.height()` after environment reset, I saw that the height is initialized as a negative value always. Therefore it is impossible for this code fragment `rewards.tolerance(physics.height(), (_STAND_HEIGHT, 2))` to return 1, since _STAND_HEIGHT is defined as 0.6 for Hopper.

Beitragsleitfaden

Beitragsleitfaden öffnen

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.