allenai / allenai/RL4LMs

Is PPO really better than SFT (in general)? under the condition of same amount of data

未關閉
#66 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
2.4k
分支
201
PR 合併指標
30 天內沒有已合併 PR

描述

For example, if we ask the model to generate a program, rather than simply continuation.

If we do not fine-tune them, RL does not even know what to generate I believe.

Do you have more thoughts on this?

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。