AlibabaResearch / AlibabaResearch/DAMO-ConvAI
Inquiry About Code Release for "Fine-Tuning Language Models with Reward Learning on Policy
Abierto
- Lenguaje dominante
- Python
- Estrellas
- 1.6k
- Forks
- 250
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
I am very interested in your paper "Fine-Tuning Language Models with Reward Learning on Policy." Could you please let me know when you plan to release the code for this work?
Thank you very much!
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Evaluación
Este issue todavía no se ha evaluado.