Is PPO really better than SFT (in general)? under the condition of same amount of data
Offen
- Vorherrschende Sprache
- Python
- Sterne
- 2.4k
- Forks
- 201
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
For example, if we ask the model to generate a program, rather than simply continuation.
If we do not fine-tune them, RL does not even know what to generate I believe.
Do you have more thoughts on this?
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Bewertung
Dieses Issue wurde noch nicht bewertet.