google-deepmind / google-deepmind/deepmind-research
Option Keyboard - GPE/GPI Experiments: FastRL Storing Q-values
Open
- Dominant language
- Jupyter Notebook
- Stars
- 15.2k
- Forks
- 2.9k
- PR merge metrics
- No merged PRs in 30d
Description
Are the q-values for different directions on the grid (q-values corresponding to up, down, left, right) from which policies are derived for fast reinforcement learning stored anywhere? I'd like to be able to generate a csv or data frame of the q-values as a function of episode number. I've looked through the code and can't find a reference to the q-values anywhere.
Contributor guide
Assessment
This issue has not been assessed yet.