Is PPO really better than SFT (in general)? under the condition of same amount of data
Aberta
- Linguagem predominante
- Python
- Estrelas
- 2.4k
- Forks
- 201
- Métricas de merge de PRs
- Nenhum PR com merge em 30d
Descrição
For example, if we ask the model to generate a program, rather than simply continuation.
If we do not fine-tune them, RL does not even know what to generate I believe.
Do you have more thoughts on this?
Guia de contribuição
Nenhum guia de contribuição indexado para este repositório
Avaliação
Esta issue ainda não foi avaliada.