Is PPO really better than SFT (in general)? under the condition of same amount of data
未關閉
- 主要語言
- Python
- 星號
- 2.4k
- 分支
- 201
- PR 合併指標
- 30 天內沒有已合併 PR
描述
For example, if we ask the model to generate a program, rather than simply continuation.
If we do not fine-tune them, RL does not even know what to generate I believe.
Do you have more thoughts on this?
貢獻指南
這個儲存庫沒有索引到貢獻指南
評估
這個 Issue 還沒有評估資料。