DLR-RM / DLR-RM/stable-baselines3
[Feature Request] independently configurable learning rates for actor and critic
Open
enhancement
help wanted
- Dominant language
- Python
- Stars
- 13.8k
- Forks
- 2.2k
- Avg merge
- 1h 35m
- Merged PRs (30d)
- 2
Description
### 🚀 Feature
independently configurable learning rates for actor and critic in AC-style algorithms
### Motivation
In literature the actor is often configured to learn slower, such that the critics responses are more reliable. At least it would be nice if i could allow my hyperparameter optimizer to decide which learning rates he wants to use for actor or critic.
### Pitch
https://github.com/DLR-RM/stable-baselines3/blob/65100a4b040201035487363a396b84ea721eb027/stable_baselines3/ddpg/ddpg.py#L12-L26
### Additional context
https://spinningup.openai.com/en/latest/algorithms/ddpg.html#documentation-pytorch-version
Contributor guide
Assessment
This issue has not been assessed yet.