AI4Finance-Foundation / AI4Finance-Foundation/ElegantRL
A policy update bug in AgentPPO?
- Lenguaje dominante
- Python
- Estrellas
- 4.4k
- Forks
- 978
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
The following codes show that the policy used to explore the env (generate the action and logprob) is 'self.act',
```
get_action = self.act.get_action
convert = self.act.convert_action_for_env
for i in range(horizon_len):
state = torch.as_tensor(ary_state, dtype=torch.float32, device=self.device)
action, logprob = [t.squeeze() for t in get_action(state.unsqueeze(0))]
```
while in the update function, the actions and policy used to calculate the 'new_log_prob' are exactly the same as the ones above:
```
new_logprob, obj_entropy = self.act.get_logprob_entropy(state, action)
ratio = (new_logprob - logprob.detach()).exp()
```
I think that 'ratio' will be always 1.
Is it a bug or there is something I misunderstand?
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Evaluación
Este issue todavía no se ha evaluado.