allenai / allenai/RL4LMs

Implementing self-play

Abierto
#18 3 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
2.4k
Forks
201
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

Hello

I would like to implement self-play dialogue training.
For that I guess I need to modify episode rollout process by adding formatting like speaker id on the start of each line. I'd also like to try holding some model buffer of previous checkpoints and use them as one of the conversants to avoid model overfitting to itself.

The obvious place for it is implementing a new policy that provides formatted generation results and holds previous checkpoints in the buffer.

Is there any better place to implement this? Anything I should consider library-wise while implementing it?

Any advice would be appreciated, thanks in advance!

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.