carpedm20 / carpedm20/deep-rl-tensorflow
Setting's of the Corridor game
Open
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 394
- PR merge metrics
- No merged PRs in 30d
Description
Could you please tell me how did you set the reward at each state? It seems that all F states will receive an reward thus an agent might just keep staying on F states till episode ends and it will automatically receive max reward. I cannot reproduce the result of the dueling network's corridor game. Could you please give me any hints?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.