AlibabaResearch / AlibabaResearch/DAMO-ConvAI
Inquiry About Code Release for "Fine-Tuning Language Models with Reward Learning on Policy
Open
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 250
- PR merge metrics
- No merged PRs in 30d
Description
I am very interested in your paper "Fine-Tuning Language Models with Reward Learning on Policy." Could you please let me know when you plan to release the code for this work?
Thank you very much!
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.