facebookresearch / facebookresearch/ReAgent

CPE Functionaility - Purpose of additional CPE Q-Networks?

Open
#456 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.7k
Forks
529
PR merge metrics
No merged PRs in 30d

Description

Hi,

I have been reading through the code and I have a few questions regarding the CPE functionality.

In particular, with the DQN model you have separate Q-networks, both normal and target, dedicated just for CPE that I would like to better understand. At present their purpose is not really clear to me. In particular, what is the purpose of these additional networks over and above the standard Q-networks of the DQN model?

In the function [_calculate_cpes](https://github.com/facebookresearch/ReAgent/blob/6c551e938f04f53c6436b700914f6212f26cccea/reagent/training/rl_trainer_pytorch.py#L216), which is part of the `RLTrainer` class. Reading this function it seems that update the networks `q_network_cpe` and `q_network_cpe_target` to model not only the reward, but any additional metrics that could be on interest in CPE.

Am I right in thinking that performing CPE on these additional metrics is the main reason for these additional networks? Put another way, if one were only interested in performing CPE on the reward itself, would using the standard Q-networks of the DQN model suffice?

Thanks

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.