Extracting total-loss, PPO-loss, rewards per step, returns per step in RLHF-PPO implementation
Open
documentation
- Dominant language
- Python
- Stars
- 4.8k
- Forks
- 487
- PR merge metrics
- No merged PRs in 30d
Description
### 📚 The doc issue
I need help extracting total-loss, PPO-loss, rewards per step, returns per step in RLHF-PPO implementation.
### Suggest a potential alternative/fix
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.