AI4Finance-Foundation / AI4Finance-Foundation/ElegantRL

maybe a small bug in the function `explore_vec_env` of discretePPO and discreteA2C?

Offen
#340 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
4.4k
Forks
978
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

in the function `explore_vec_env` of `AgentPPO`, the variable `actions` shaped with `[horizon_len, self.num_envs, 1]`, but the following expression `convert(action)` return the tensor with the 1-dim shape `num_envs`, which actually should be `[num_envs, 1]` as it works in `explore_vec_env` of `AgentD3QN`. And it indeed faild the demo`examples/demo_A2C_PPO.py`.

Folloiwing change works for me:
```python
# ActorDiscretePPO of net.py
def get_action(self, state: Tensor) -> (Tensor, Tensor):
state = self.state_norm(state)
a_prob = self.soft_max(self.net(state))
a_dist = self.ActionDist(a_prob)
action = a_dist.sample()
logprob = a_dist.log_prob(action)
return action.unsqueeze(1), logprob # unsqueeze the action
```

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.